Papers with Natural Language Interfaces

300 papers
One Time of Interaction May Not Be Enough: Go Deep with an Interaction-over-Interaction Network for Response Selection in Dialogues (P19-1)

Copied to clipboard

Challenge: Currently, retrieval-based dialogues are performed in shallow ways . a recent study investigated the problem of context-response matching in open-domain .
Approach: They propose a model that lets utterance-response interaction go deep by stacking interaction blocks.
Outcome: The proposed model outperforms state-of-the-art methods on three benchmark data sets.
SF-QA: Simple and Fair Evaluation Library for Open-domain Question Answering (2021.eacl-demos)

Copied to clipboard

Challenge: Open-domain question answering (QA) requires large amounts of resources and is difficult to reproduce results due to complex configurations.
Approach: They propose a simple and fair evaluation framework for open-domain question answering (QA) it modularizes the pipeline open- domain QA system, making it easily accessible .
Outcome: The proposed evaluation framework is publicly available and anyone can contribute to the code and evaluations.
PropGenie: A Multi-Agent Conversational Framework for Real Estate Assistance (2026.eacl-demo)

Copied to clipboard

Challenge: PropGenie is a multi-agent framework based on large language models (LLMs) it provides comprehensive real estate assistance in real-world scenarios .
Approach: They propose a multi-agent framework based on large language models to deliver comprehensive real estate assistance in real-world scenarios.
Outcome: The proposed framework outperforms a general-purpose LLM and a domain-specific chatbot in real-world scenarios.
NAMER: A Node-Based Multitasking Framework for Multi-Hop Knowledge Base Question Answering (2021.naacl-demos)

Copied to clipboard

Challenge: Using a node-based framework, knowledge base question answering systems can grasp structural mappings between questions and KB queries.
Approach: They propose a node-based framework that better grasps the structural mapping between questions and KB queries by aligning the nodes in a query with their corresponding mentions in question.
Outcome: The proposed framework outperforms the previous SoTA on CCKS CKBQA dataset.
SpiritRAG: A Q&A System for Religion and Spirituality in the United Nations Archive (2025.emnlp-demos)

Copied to clipboard

Challenge: Religion and spirituality (R/S) are complex and domain-dependent concepts that have long confounded researchers and policymakers.
Approach: They propose an interactive question-answering system based on Retrieval-Augmented Generation (RAG) SpiritRAG allows researchers and policymakers to conduct complex, context-sensitive database searches of large datasets .
Outcome: SpiritRAG is an interactive Q&A system based on Retrieval-Augmented Generation (RAG) built using 7,500 UN resolution documents related to religion and spirituality in the domains of health and education.
Summarization of Dialogues and Conversations At Scale (2023.eacl-tutorials)

Copied to clipboard

Challenge: Conversations are the natural communication format for people.
Approach: This tutorial will survey the cutting-edge methods for summarizing written and spoken conversation.
Outcome: This tutorial will examine the cutting-edge methods for summarizing written and spoken conversations, covering key sub-areas whose combination is needed for a successful solution.
Fast Prototyping a Dialogue Comprehension System for Nurse-Patient Conversations on Symptom Monitoring (N19-2)

Copied to clipboard

Challenge: a limited amount of data exists for human-human spoken dialogues for research and development . a dialogue comprehension system that extracts clinical information from spoken conversations is clinically useful .
Approach: They propose a framework inspired by nurse-initiated clinical symptom monitoring conversations to construct a simulated human-human dialogue dataset.
Outcome: The proposed system achieves more than 80% F1 on held-out test set from nurse-to-patient conversations.
Optimal Summaries for Enabling a Smooth Handover in Chat-Oriented Dialogue (2022.aacl-srw)

Copied to clipboard

Challenge: In dialogue systems, it is difficult to provide fully autonomous dialogue . to ensure a good dialogue experience, human operators sometimes need to intervene .
Approach: They conducted large-scale experiments on chat dialogues to determine which type of summary is most useful for handover . abstractive summary plus one utterance immediately before handover and extractive summary consisting of five utterrances immediately before the handover were found to be the most useful .
Outcome: The best summaries were abstractive summary plus one utterance before handover and extractive summary consisting of five utterrances before hand over.
Augmenting Transformers with KNN-Based Composite Memory for Dialog (2021.tacl-1)

Copied to clipboard

Challenge: Recent work has focused on learning architectures with large memories capable of storing external knowledge.
Approach: They propose a method to augment generative Transformer neural networks with information fetching modules.
Outcome: The proposed approach improves performance in generative dialog modeling . external knowledge is retrieved from Wikipedia, images, and human-written dialog utterances .
Deep Learning for Conversational AI (N18-6)

Copied to clipboard

Challenge: Spoken Dialogue Systems (SDS) have great commercial potential . the advent of deep learning has led to significant advances in this area of NLP research .
Approach: This tutorial will introduce researchers to the pipeline framework for modelling goal-oriented dialogue systems.
Outcome: This tutorial will familiarise researchers with the latest advances in spoken dialogue systems . the aim of the course is to encourage dialogue research in the NLP community .
Best Practices for Data-Efficient Modeling in NLG:How to Train Production-Ready Neural Models with Less Data (2020.coling-industry)

Copied to clipboard

Challenge: Natural language generation (NLG) is a critical component in conversational systems . Traditionally, NLG components have been deployed using template-based solutions . however, deployment of such model-based systems has been challenging due to high latency and data needs.
Approach: They propose a family of techniques to deploy data-efficient neural solutions for NLG in conversational systems to production.
Outcome: The proposed techniques achieve production quality with light-weight neural network models using fraction of the data needed otherwise.
Annotation Process for the Dialog Act Classification of a Taglish E-commerce Q&A Corpus (D19-51)

Copied to clipboard

Challenge: Existing studies on DA classification in general contexts have not addressed this problem.
Approach: They constructed a text-based corpus of 7,265 posts from the question and answer section of products on Lazada Philippines.
Outcome: The text-based corpus of 7,265 posts from the question and answer section of products on Lazada Philippines was constructed using a tagset for DA classification . the corpus was composed dominantly of single-label posts, with 34% of the corpuse having multiple intent tags.
Towards Answer-unaware Conversational Question Generation (D19-58)

Copied to clipboard

Challenge: Existing frameworks for conversational question generation are answeraware, but are not able to generate corresponding answers . a number of question generation methods are developed for text-based question answering .
Approach: They propose a framework for conversational question generation that is unaware of the corresponding answers.
Outcome: The proposed framework is effective but answeraware, the authors show . it improves quality of generated questions if question foci and question patterns are identified .
Extracting relevant information from physician-patient dialogues for automated clinical note taking (D19-62)

Copied to clipboard

Challenge: a system that extracts pertinent medical information from dialogues between clinicians and patients is proposed . entering data into EMRs is currently slow and error-prone, and clinicians spend up to 50% of their time on data entry.
Approach: They propose a system that automatically extracts medical information from dialogues between clinicians and patients using context and time information.
Outcome: The proposed system extracts medical information from dialogues and automatically generates a patient note.
Is He Extroverted? Identifying Missing Relevant Personas for Faithful User Simulation (2026.eacl-srw)

Copied to clipboard

Challenge: Existing user simulation approaches focus on generating user-like responses in dialogue without verifying whether critical personas are supplied.
Approach: They propose a task of identifying persona dimensions that are relevant but missing in simulating a user's reply for a given dialogue context.
Outcome: The proposed model identifies persona dimensions that are relevant but missing in simulating a user’s response for a given dialogue context.
Dynamic Schema Graph Fusion Network for Multi-Domain Dialogue State Tracking (2022.acl-long)

Copied to clipboard

Challenge: Existing approaches to model the relations between domains and slots fail to address these issues and can be generalized to unseen domains.
Approach: They propose a Dynamic Schema Graph Fusion Network which generates a dynamic schema graph to explicitly fuse prior slot-domain membership relations and dialogue-aware dynamic slot relations.
Outcome: The proposed model outperforms existing methods on benchmark datasets showing that it can extract users' goals or intentions as dialogue states and keep them updated over the whole dialogue.
Slot-consistent NLG for Task-oriented Dialogue Systems with Iterative Rectification Network (2020.acl-main)

Copied to clipboard

Challenge: Existing approaches to natural language generation are prone to errors, such as neglecting input slot values and generating redundant slot values.
Approach: They propose an iterative rectification network to improve general NLG systems . they apply bootstrapping algorithms to sample training candidates and incorporate reward .
Outcome: The proposed methods significantly reduce the slot error rate for strong baselines.
Evaluating and Modeling Attribution for Cross-Lingual Question Answering (2023.emnlp-main)

Copied to clipboard

Challenge: Open-retrieval question answering systems are lacking in attribution for cross-lingual question answering . open-research questions are available in 20 languages, but their raw generation often falls short in factuality .
Approach: They are the first to study attribution for cross-lingual question answering . they collect data in 5 languages to assess the attribution level of a state-of-the-art QA system .
Outcome: The proposed approach improves the attribution level of a state-of-the-art cross-lingual QA system.
TruthReader: Towards Trustworthy Document Assistant Chatbot with Reliable Attribution (2024.emnlp-demo)

Copied to clipboard

Challenge: Document assistant chatbots are empowered with extensive capabilities by Large Language Models (LLMs) however, they suffer from hallucinations that are difficult to verify in the context of given documents.
Approach: They propose a document assistant chatbot with reliable attribution that enables users to seek relevant information from given documents.
Outcome: The proposed system generates answers with detailed inline citations, which can be attributed to the original document paragraphs, facilitating verification of factual consistency of the generated text.
Knowledge-centered conversational agents with a drive to learn (2024.naacl-srw)

Copied to clipboard

Challenge: Unlike traditional task-oriented dialogue agents, knowledgeable agents can autonomously determine what they know and do not know, what is the epistemic status of what they do not understand, and what they need to learn.
Approach: They propose an adaptive conversational agent that assesses the quality of its knowledge and is driven to become more knowledgeable.
Outcome: The proposed agent can learn effective policies to acquire the knowledge needed by assessing the efficiency of these capabilities during interaction.
SceMQA: A Scientific College Entrance Level Multimodal Question Answering Benchmark (2024.acl-short)

Copied to clipboard

Challenge: SceMQA focuses on core science subjects including Mathematics, Physics, Chemistry, and Biology.
Approach: They propose to use SceMQA to evaluate multimodal question answering at college entrance level.
Outcome: The proposed model provides specific knowledge points for each problem and detailed explanations for each answer.
Multi-hop Selector Network for Multi-turn Response Selection in Retrieval-based Chatbots (D19-1)

Copied to clipboard

Challenge: Existing studies focus on matching candidate responses with every context utterance, but it also brings noise signals and unnecessary information.
Approach: They propose a multi-hop selector network to match context with candidate responses . they propose to use a selector to filter the relevant utterances as context .
Outcome: The proposed model outperforms state-of-the-art methods on three public multi-turn dialogue datasets.
MoEL: Mixture of Empathetic Listeners (D19-1)

Copied to clipboard

Challenge: Neural network approaches for conversation models have shown to be successful in generating fluent and relevant responses.
Approach: They propose a novel end-to-end approach for modeling empathy in dialogue systems by using Mixture of Empathetic Listeners (MoEL).
Outcome: The proposed model outperforms multitask training baseline in terms of empathy, relevance, and fluency.
Human-Like Embodied AI Interviewer: Employing Android ERICA in Real International Conference (2025.coling-demos)

Copied to clipboard

Challenge: Qualitative interviews are foundational to social science research, offering deep insights through open-ended conversations.
Approach: They introduce a human-like embodied AI interviewer which integrates android and humanoid robots equipped with advanced conversational capabilities.
Outcome: The proposed system performs well in a real-world case study at SIGDIAL 2024 with 42 participants, of whom 69% reported positive experiences.
Are the Tools up to the Task? an Evaluation of Commercial Dialog Tools in Developing Conversational Enterprise-grade Dialog Systems (N19-2)

Copied to clipboard

Challenge: Existing toolsets are incomplete in meeting the goal of building effective dialog systems, authors say .
Approach: They compare dialog tools available from a number of companies to determine their strengths and weaknesses . they provide quantitative and qualitative results in three main areas: natural language understanding, dialog, and text generation .
Outcome: The toolsets are incomplete, but they are compared to other tools to determine their strengths and weaknesses.
Athena 2.0: Contextualized Dialogue Management for an Alexa Prize SocialBot (2021.emnlp-demo)

Copied to clipboard

Challenge: Athena 2.0 is a socialbot that has been a finalist in the last two Alexa Prize Grand Challenges.
Approach: They describe Athena 2.0's dialogue management strategy and its performance in the Alexa Prize 20/21 competition.
Outcome: The system is a finalist in the Alexa Prize 20/21 competition and will be shown on a live demo and recorded video recordings.
ADVISER: A Dialog System Framework for Education & Research (P19-3)

Copied to clipboard

Challenge: In this paper, we focus on task-oriented dialog systems, although our framework allows easy integration of non-task dialog systems and their combination.
Approach: They propose an open source dialog system framework for education and research that supports multi-domain task-oriented conversations in two languages.
Outcome: The proposed framework supports multi-domain task-oriented conversations in two languages and is open source for education and research.
Asking the Right Question at the Right Time: Human and Model Uncertainty Guidance to Ask Clarification Questions (2024.eacl-long)

Copied to clipboard

Challenge: Using model uncertainty as supervision for deciding when to ask may not be the most effective way to resolve model uncertainty.
Approach: They propose to generate clarification questions based on model uncertainty estimation and compare it to several alternatives to generate questions .
Outcome: The proposed approach improves the model uncertainty of a collaborative dialogue task and shows that it is more effective than other alternatives.
ScoutBot: A Dialogue System for Collaborative Navigation (P18-4)

Copied to clipboard

Challenge: Demo will allow users to issue unconstrained spoken language commands to ScoutBot.
Approach: The demonstration will allow users to issue unconstrained spoken language commands to ScoutBot.
Outcome: The demonstration will allow users to issue unconstrained spoken language commands to ScoutBot.
Development of Conversational AI for Sleep Coaching Programme (2021.eacl-srw)

Copied to clipboard

Challenge: Existing methods to treat insomnia neglect conversational aspects, which plays a critical role in sleep therapy.
Approach: They propose to develop conversational AI for a sleep coaching programme which is motivated by CBT-I treatment and provide an automated analytic system to support human experts.
Outcome: The proposed system could interact naturally with a user and provide an automated analytic system to support human experts.
An Emotional Comfort Framework for Improving User Satisfaction in E-Commerce Customer Service Chatbots (2021.naacl-industry)

Copied to clipboard

Challenge: E-commerce has grown rapidly over the last several years, and chatbots for intelligent customer service are simultaneously drawing attention.
Approach: They propose a framework to obtain proper answer to customers’ emotional questions using emotion classification model and text matching.
Outcome: The proposed framework is very promising on real online systems.
PEEP-Talk: A Situational Dialogue-based Chatbot for English Education (2023.acl-demo)

Copied to clipboard

Challenge: Existing chatbots lack realistic practice scenarios for English learners . existing platforms employ hand-crafted and patternmatching rules, limiting communication ability and responding appropriately to out-of-situation utterances.
Approach: They propose a real-world situational dialogue-based chatbot for English education . it generates appropriate responses in various real-life situations while providing accurate feedback .
Outcome: The proposed chatbot generates appropriate responses in various real-life situations while providing accurate feedback to learners.
Leveraging Large Language Models for Conversational Multi-Doc Question Answering: The First Place of WSDM Cup 2024 (2025.findings-acl)

Copied to clipboard

Challenge: WSDM Cup 2024 presents a challenge for conversational multi-doc question answering using large language models . a hybrid training strategy is developed to make the most of in-domain unlabeled data .
Approach: They propose a conversational multi-doc question answering challenge in WSDM Cup 2024 . they adapt LLMs to the task, then devise a hybrid training strategy to make the most of unlabeled data.
Outcome: The proposed approach ranked 1st in the WSDM Cup 2024 challenge . it exploits the superior natural language understanding and generation capability of Large Language Models .
Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Question Answering Evaluation (2022.emnlp-main)

Copied to clipboard

Challenge: despite the importance of question answering, evaluations of QA systems are typically limited by manual annotations . despite this, little progress has been made in QA evaluations based on a single answer .
Approach: They propose to extend over exact match (EM) with predefined rules or token-level F1 measure . they propose to use a BERT matching measure to approximate QA predictions .
Outcome: The proposed model improves AE approximations and more accurately reflects the performance of systems.
ClinQueryAgent: A Conversational Agent for Population Health Management (2026.acl-demo)

Copied to clipboard

Challenge: In 2014 there were 2.3B SNOMED CT codes recorded in English healthcare practices. By 2024, this number had grown almost threefold to 6.1B codes: approximately 100 codes per person each year.
Approach: They introduce a system for translating natural language population health questions into executable database queries using agents with access to both local and external knowledge bases.
Outcome: The proposed system is able to handle a range of health informatics tasks on three datasets and via a beta-testing phase.
Interactive Instance-based Evaluation of Knowledge Base Question Answering (D18-2)

Copied to clipboard

Challenge: Existing approaches to Knowledge Base Question Answering are based on semantic parsing.
Approach: They propose a tool that aids in debugging of question answering systems that construct a structured semantic representation for the input question.
Outcome: The proposed system allows debugging of model predictions on individual instances and simplifies manual error analysis.
Tab-CQA: A Tabular Conversational Question Answering Dataset on Financial Reports (2023.acl-industry)

Copied to clipboard

Challenge: Existing conversational question answering datasets are usually constructed from unstructured texts in English.
Approach: They propose a Chinese tabular conversational question answering dataset based on financial reports . they select 2,463 tables and manually generate 2,463, conversations with 35,494 QA pairs .
Outcome: The proposed dataset is based on Chinese financial reports extracted from listed companies in the past 30 years.
Did Aristotle Use a Laptop? A Question Answering Benchmark with Implicit Reasoning Strategies (2021.tacl-1)

Copied to clipboard

Challenge: Existing questions that explicitly describe the process for deriving the answer are often implicit.
Approach: They propose a question answering benchmark where the required reasoning steps are implicit in the question and should be inferred using a strategy.
Outcome: The proposed model is short, topic-diverse, and covers a wide range of strategies.
DeepPavlov: Open-Source Library for Dialogue Systems (P18-4)

Copied to clipboard

Challenge: open-source library DeepPavlov is designed for rapid development of dialogue systems.
Approach: open-source library DeepPavlov is tailored for development of conversational agents . the library prioritizes efficiency, modularity and extensibility with the goal to make it easier to develop dialogue systems from scratch .
Outcome: the open-source library DeepPavlov is designed for rapid development of dialogue systems . it supports modular as well as end-to-end approaches to implementation of conversational agents .
PeerQA: A Scientific Question Answering Dataset from Peer Reviews (2025.naacl-long)

Copied to clipboard

Challenge: a dataset of 579 QA pairs from 208 scientific articles contains answers that reviewers raised while thoroughly examining the scientific article.
Approach: They propose a dataset that contains questions that reviewers raised while thoroughly examining the scientific article.
Outcome: The proposed dataset contains 579 QA pairs from 208 academic articles . the results show that decontextualization approaches improve retrieval performance .
Structural Characterization for Dialogue Disentanglement (2022.acl-long)

Copied to clipboard

Challenge: tangled multi-party dialogues lead to difficulties in understanding the dialogue history for both human and machine.
Approach: They propose a model for disentangling multi-party dialogues using speaker property and reference dependency.
Outcome: The proposed model achieves state-of-the-art on the Ubuntu IRC benchmark dataset and contributes to dialogue-related comprehension.
Recipes for Building an Open-Domain Chatbot (2021.eacl-main)

Copied to clipboard

Challenge: Existing work shows that scaling models in the number of parameters and the size of the data they are trained on gives improved results, but other factors are important.
Approach: They propose to build open-domain chatbots that can be scaled to improve their performance . they use a blend of cognitive and cognitive skills to build a model that combines these skills .
Outcome: The proposed models outperform existing approaches in multi-turn dialogue on engagingness and humanness measurements.
Interactive Task Learning from GUI-Grounded Natural Language Instructions and Demonstrations (2020.acl-demos)

Copied to clipboard

Challenge: SUGILITE is an intelligent task automation agent that can learn new tasks and relevant associated concepts interactively from the user’s natural language instructions and demonstrations using GUIs.
Approach: They propose to use third-party mobile apps to teach new tasks and concepts using verbal instructions and demonstrations.
Outcome: The proposed system can learn new tasks and relevant concepts from user's natural language instructions and demonstrations, and it generalizes taught concepts to different contexts and task domains.
UFO: A UI-Focused Agent for Windows OS Interaction (2025.naacl-long)

Copied to clipboard

Challenge: UFO is a UI-Fcused agent designed to fulfill user requests tailored to Windows OS applications . it decomposes user requests using divide-and-conquer approach, enabling seamless navigation and addressing sub-tasks across multiple applications.
Approach: They propose a UI-Fcused Windows OS agent that decomposes user requests using a divide-and-conquer approach and incorporates a control interaction module tailored for Windows OS.
Outcome: The proposed agent decomposes user requests using divide-and-conquer approach, enabling seamless navigation and addressing sub-tasks across multiple applications.
Samvaadhana: A Telugu Dialogue System in Hospital Domain (D19-61)

Copied to clipboard

Challenge: a dialogue system for Hospital domain in Telugu is a resource-poor Dravidian language . the system handles various hospital and doctor related queries .
Approach: They propose to model a dialogue system for Hospital domain in Telugu which is a resource-poor Dravidian language.
Outcome: The proposed system achieves a high overall rating and a significantly accurate context-capturing method.
The Why and The How: A Survey on Natural Language Interaction in Visualization (2022.naacl-main)

Copied to clipboard

Challenge: Recent research shows that different forms of natural language-based interaction prove suitable to support users in accomplishing various visualization tasks.
Approach: They propose a taxonomy of visualization tasks and a classification system to illustrate the state-of-the-art of natural language-based interaction in visualization.
Outcome: The proposed model can support annotations, recommendations, explanations, and documentation tasks.
InSCIt: Information-Seeking Conversations with Mixed-Initiative Interactions (2023.tacl-1)

Copied to clipboard

Challenge: In information-seeking conversations, a user may ask questions that are under-specified or unanswerable.
Approach: They present a dataset for information-seeking conversations with mixed-initiative interactions . they use Wikipedia to search for answers and provide relevant information .
Outcome: The proposed system significantly underperforms humans in two of the most recent studies.
Cross-Lingual Open-Domain Question Answering with Answer Sentence Generation (2022.aacl-main)

Copied to clipboard

Challenge: Open-Domain Generative Question Answering has achieved impressive performance in English . combining document-level retrieval with answer generation can generate complete sentences .
Approach: They propose an open-domain approach that combines document retrieval with answer generation to generate complete sentences in English . they propose a cross-lingual generative model that exploits passages written in multiple languages .
Outcome: The proposed model outperforms answer sentence selection baselines for all 5 languages and monolingual pipelines for three out of five languages.
Action and Reaction Go Hand in Hand! a Multi-modal Dialogue Act Aided Sarcasm Identification (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies have shown that sarcasm is reflected by the intended meaning of the speaker's utterance.
Approach: They propose to extend the MUStARD dataset to enclose dialogue acts for each dialogue . they propose a dialogue act-aided multi-modal transformer network for sarcasm identification model .
Outcome: The proposed model improves performance in dialogue act-aided sarcasm identification compared to sardasmatic identification alone.
Towards Transparent Interactive Semantic Parsing via Step-by-Step Correction (2022.findings-acl)

Copied to clipboard

Challenge: Existing studies on semantic parsing focus on mapping a natural-language utterance to a logical form (LF) but natural language may contain ambiguity and variability, making this challenge difficult.
Approach: They propose an interactive semantic parsing framework that explains the predicted LF step by step in natural language and enables the user to make corrections through natural-language feedback for individual steps.
Outcome: The proposed framework improves parsing accuracy and transparency in a crowdsourced dialogue dataset.
Semantic XPath: Structured Agentic Memory Access for Conversational AI (2026.acl-demo)

Copied to clipboard

Challenge: Early ConvAI agents rely on an in-context approach that appends the growing conversation history to the model input, but this approach scales poorly under context-window limits.
Approach: They propose a tree-structured memory module to access and update structured conversational memory.
Outcome: The proposed system improves over flat-RAG baselines while using only 9.1% of the tokens required by in-context memory.
Developing a Production System for Purpose of Call Detection in Business Phone Conversations (2022.naacl-industry)

Copied to clipboard

Challenge: a commercial system detects Purpose of Call statements in call transcripts . the model is based on a set of rules and a neural model .
Approach: They propose a system to detect Purpose of Call statements in English business call transcripts in real time.
Outcome: The proposed model achieves 88.6 F1 on average in various types of business calls and has low inference time.
UHop: An Unrestricted-Hop Relation Extraction Framework for Knowledge-Based Question Answering (N19-1)

Copied to clipboard

Challenge: Existing work restricts search from one entity to another to the maximum number of hops . a knowledge graph is a powerful graph structure that encodes knowledge to save and organize it .
Approach: They propose an unrestricted-hop framework which relaxes the restriction by using a transition-based search framework.
Outcome: The proposed framework performs well with state-of-the-art models and is competitive without exhaustive searches.
CR-GIS: Improving Conversational Recommendation via Goal-aware Interest Sequence Modeling (2022.coling-1)

Copied to clipboard

Challenge: Existing methods to determine a goal item by sequentially tracking users’ interests ignore the rich goal-aware implicit interest sequence patterns in a dialog.
Approach: They propose to model goal-aware implicit user interest sequence patterns in a dialog and a hierarchical Star Transformer to guide multi-turn utterances generation.
Outcome: The proposed framework achieves more accurate recommendations with more fluent and coherent dialog utterances.
Evaluating Open-Domain Dialogues in Latent Space with Next Sentence Prediction and Mutual Information (2023.acl-long)

Copied to clipboard

Challenge: Existing evaluation methods for open-domain dialogues are difficult due to the one-to-many issue of the open- domain dialogues.
Approach: They propose a learning-based automatic evaluation metric which can robustly evaluate open-domain dialogues by augmenting CVAEs with a Next Sentence Prediction objective and employing Mutual Information to model the semantic similarity of text in the latent space.
Outcome: The proposed method can evaluate open-domain dialogues on two open- domain dialogue datasets.
I already said that! Degenerating redundant questions in open-domain dialogue systems. (2023.acl-srw)

Copied to clipboard

Challenge: Neural text generation models have been successful in short open-domain conversations, but their performance degrades significantly in the long term.
Approach: They propose a method to generate training data without crowdsourcing . they adapt negative training, decoding, and classification methods to mitigate redundancy problem .
Outcome: The proposed method reduces the rate of redundant questions from 27.2% to 8.7% while improving the quality of the original model.
Reinforced Question Rewriting for Conversational Question Answering (2022.emnlp-industry)

Copied to clipboard

Challenge: Existing approaches to CQA involve training new models from scratch . existing approaches are expensive and often not feasible .
Approach: They propose to use QA feedback to supervise the rewriting model with reinforcement learning.
Outcome: The proposed model can improve QA performance over baselines for extractive and retrieval QA.
Which is Better for Deep Learning: Python or MATLAB? Answering Comparative Questions in Natural Language (2021.eacl-demos)

Copied to clipboard

Challenge: Comparative QA is a challenging task since it requires collecting evidence from many different sources.
Approach: They propose a natural language interface for comparative QA that can be used in personal assistants, chatbots, and similar NLP devices.
Outcome: The proposed system can be used in personal assistants, chatbots, and similar NLP devices.
CRWIZ: A Framework for Crowdsourcing Real-Time Wizard-of-Oz Dialogues (2020.lrec-1)

Copied to clipboard

Challenge: Crowdsourcing platforms such as Amazon Mechanical Turk have been effective for collecting large corpora of task-based and open-domain conversational dialogues, but difficulties arise when task- based dialogues require expert domain knowledge or rapid access to domain-relevant information.
Approach: They propose a framework for collecting real-time Wizard of Oz dialogues through crowdsourcing for collaborative, complex tasks.
Outcome: The proposed framework avoids interactions that breach procedures only known to experts while enabling the capture of a wide variety of interactions.
Hybrid Graphs for Table-and-Text based Question Answering using LLMs (2025.naacl-long)

Copied to clipboard

Challenge: Current methods for QA rely on fine-tuning and high-quality data, which is difficult to obtain.
Approach: They propose a Hybrid Graph-based approach for Table-Text QA that leverages Large Language Models without fine-tuning.
Outcome: The proposed approach improves Exact Match scores by 10% on Hybrid-QA and 5.4% on OTT-QA.
Open Domain Question Answering over Tables via Dense Retrieval (2021.naacl-main)

Copied to clipboard

Challenge: Recent advances in open-domain QA focus on retrieving textual passages . a retriever designed to handle tabular context can improve retrieval quality .
Approach: They propose a tabular-based retrieval model that improves retrieval quality over a BERT-based retriever.
Outcome: The proposed retriever improves retrieval quality with mined hard negatives over a BERT-based retriever.
DebateQA: Evaluating Question Answering on Debatable Knowledge (2026.findings-eacl)

Copied to clipboard

Challenge: Existing QA benchmarks that provide fixed answers to debatable questions are inadequate for evaluating their performance.
Approach: They propose to use a dataset of 2,941 debatable questions to assess their ability to provide comprehensive answers to inherently debatably asked questions.
Outcome: The proposed model performs well on 2,941 debatable questions accompanied by human-annotated partial answers that capture a variety of perspectives.
Toward Implicit Reference in Dialog: A Survey of Methods and Data (2022.aacl-main)

Copied to clipboard

Challenge: In natural language, speakers often leave out information that is understood by the other party through the shared context.
Approach: They propose to use omitted entities as implicit references in dialogs to improve language processing.
Outcome: The proposed method is based on a set of experiments which show that the proposed method has a high level of accuracy and is a success.
Interactive Text-to-Image Retrieval with Large Language Models: A Plug-and-Play Approach (2024.acl-long)

Copied to clipboard

Challenge: primarily addressed in text-to-image retrieval task using dialogue-form context query . conventionally, text-based retrieval methods rely on initial text descriptions .
Approach: They propose a plug-based retrieval method that uses large language models as questioners to generate non-redundant questions about the attributes of the target image.
Outcome: The proposed method performs better than zero-shot and fine-tuned baselines in benchmarks.
Towards Question-Answering as an Automatic Metric for Evaluating the Content Quality of a Summary (2021.tacl-1)

Copied to clipboard

Challenge: Existing text overlap based evaluation metrics are limited to matching tokens, either lexically or via embeddings.
Approach: They propose a metric to evaluate the content quality of a summary using question-answering (QA) QA-based methods directly measure a summary’s information overlap with a reference, making them fundamentally different from text overlap metrics.
Outcome: The proposed metric outperforms current state-of-the-art metrics on most evaluations using benchmark datasets while being competitive on others due to limitations of state- of-the art models.
Towards Better Generalization in Open-Domain Question Answering by Mitigating Context Memorization (2024.findings-naacl)

Copied to clipboard

Challenge: Open-domain Question Answering (OpenQA) aims at answering factual questions using an external large-scale knowledge corpus.
Approach: They propose a retrieval-augmented approach to QA that focuses on retrieving relevant knowledge from an external corpus.
Outcome: The proposed model can generalize to completely different knowledge domains while adapting to updated versions of the same knowledge corpus and switching to completely new knowledge domain.
IndustryAssetEQA: A Neurosymbolic Operational Intelligence System for Embodied Question Answering in Industrial Asset Maintenance (2026.acl-industry)

Copied to clipboard

Challenge: Industrial maintenance assistants produce generic explanations that are weakly grounded in telemetry and omit verifiable provenance.
Approach: They propose a neurosymbolic operational intelligence system that combines episode-centric telemetry representations with a Failure Mode and Effects Analysis Knowledge Graph to enable Embodied Question Answering over industrial assets.
Outcome: The proposed system improves structural validity by up to +0.51, counterfactual accuracy by up . to +0.47, and explanation entailment by +0.64, while reducing severe expert-rated overclaims from 28% to 2%.
Chart Question Answering from Real-World Analytical Narratives (2025.acl-srw)

Copied to clipboard

Challenge: a dataset for chart question answering is constructed from visualization notebooks . data visualizations are an essential modality for communicating complex information about data.
Approach: They propose a dataset for chart question answering constructed from visualization notebooks . they use real-world, multi-view charts paired with natural language questions .
Outcome: The proposed dataset is constructed from student-authored visualization notebooks . it features real-world, multi-view charts paired with natural language questions . initial evaluations highlight significant performance gaps .
Augmenting Compliance-Guaranteed Customer Service Chatbots: Context-Aware Knowledge Expansion with Large Language Models (2025.emnlp-industry)

Copied to clipboard

Challenge: Retrieval-based chatbots leverage human-verified Q&A knowledge to deliver accurate, verifiable responses.
Approach: They propose a similar question generation task for LLM training and inference to enable comprehensive semantic exploration and enhanced alignment with source question-answer relationships.
Outcome: The proposed methods achieve 92% user satisfaction rate in a deployed chatbot system, reflecting an 18% improvement over the baseline.
It is AI’s Turn to Ask Humans a Question: Question-Answer Pair Generation for Children’s Story Books (2022.acl-long)

Copied to clipboard

Challenge: Existing question answering (QA) techniques are created mainly to answer questions asked by humans, but in educational applications, teachers often need to decide what questions to ask .
Approach: They propose to use a fairytale-themed storybook as input to generate QA pairs that can test a student's comprehension skills.
Outcome: The proposed system outperforms state-of-the-art QAG baseline systems and builds an interactive story-telling application for the future real-world deployment.
Know Better – A Clickbait Resolving Challenge (2022.lrec-1)

Copied to clipboard

Challenge: a clickbait headline or teaser is used to "bait" the reader into clicking a link to an article . clickbaiting is annoying but effective, and can be countered with specialized models .
Approach: They propose to construct approaches that can automatically extract relevant information from clickbait articles . they argue that clickbaiting can probably not be defeated with clickbaitting detection alone .
Outcome: The proposed methods outperform question answering models on clickbait resolving task . the data will be used to give users tools to counter clickbaiting in the future .
CarExpert: Leveraging Large Language Models for In-Car Conversational Question Answering (2023.emnlp-industry)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated remarkable performance by following natural language instructions without fine-tuning them on domain-specific tasks and data.
Approach: They propose an in-car retrieval-augmented conversational question-answering system that uses large language models to generate natural, safe and domain-specific answers.
Outcome: The proposed system outperforms state-of-the-art LLMs in generating safe and domain-specific answers.
Training Adaptive Computation for Open-Domain Question Answering with Computational Constraints (2021.acl-short)

Copied to clipboard

Challenge: Adaptive Computation (AC) has been shown to be effective in improving the efficiency of Open-Domain Question Answering systems.
Approach: They propose an AC method that can be applied to an existing ODQA model and can be trained efficiently on a single GPU.
Outcome: The proposed method improves upon a state-of-the-art model on two datasets and is more accurate than previous AC methods due to the stronger base ODQA model.
BLISS: An Agent for Collecting Spoken Dialogue Data about Health and Well-being (2020.lrec-1)

Copied to clipboard

Challenge: Structured interviews are a time-consuming and inefficient way to gather information about people's well-being.
Approach: They propose to build an artificial intelligence agent which asks questions about happiness . they build a prototype of the agent and collect 55 spoken dialogues .
Outcome: The proposed agent collects 55 spoken dialogues and asks users about happiness and well-being.
The JDDC Corpus: A Large-Scale Multi-Turn Chinese Dialogue Dataset for E-commerce Customer Service (2020.lrec-1)

Copied to clipboard

Challenge: Existing datasets for human-like dialogue tasks are deficient due to the complexity of human conversations.
Approach: They construct a large-scale Chinese E-commerce conversation corpus with 1 million dialogues, 20 million utterances, and 150 million words.
Outcome: The proposed dataset includes 1 million multi-turn dialogues, 20 million utterances, and 150 million words.
DialMed: A Dataset for Dialogue-based Medication Recommendation (2022.coling-1)

Copied to clipboard

Challenge: Existing studies on medication recommendation mainly rely on EHRs, but some details of interactions between doctors and patients may be ignored or omitted in EHR.
Approach: They propose to use medical dialogues to recommend medications with medical dialogue data . they propose to model dialogue structure and disease knowledge aware network .
Outcome: The proposed method is a promising solution to recommend medications with medical dialogues.
How to Motivate Your Dragon: Teaching Goal-Driven Agents to Speak and Act in Fantasy Worlds (2021.naacl-main)

Copied to clipboard

Challenge: a recent improvement in the quality of natural language processing and generation (NLG) is needed for goal-oriented ML driven agents.
Approach: They propose a reinforcement learning system that integrates large-scale language modeling and commonsense reasoning-based pre-training to imbue the agent with relevant priors.
Outcome: The proposed system is able to act and talk naturally with respect to their motivations.
Fusing Temporal Graphs into Transformers for Time-Sensitive Question Answering (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for extracting temporal information from text are not suitable for time-sensitive questions.
Approach: They propose to use existing temporal information extraction systems to construct temporal graphs of events, times, and temporal relations in questions and documents.
Outcome: The proposed method outperforms graph convolution-based approaches on SituatedQA and TimeQA.
Intent-calibrated Self-training for Answer Selection in Open-domain Dialogues (2023.tacl-1)

Copied to clipboard

Challenge: Existing answer selection models require large amounts of labeled data to produce accurate answers.
Approach: They propose intent-calibrated self-training to calibrate answer labels using labeled data . they propose intentcalibration to improve quality of pseudo answer labels .
Outcome: The proposed intent-calibrated answer selection paradigm outperforms baselines with 1%, 5%, and 10% labeled data on two benchmark datasets.
A Few More Examples May Be Worth Billions of Parameters (2022.findings-emnlp)

Copied to clipboard

Challenge: Recent work on few-shot learning for natural language tasks explores the dynamics of scaling up either the number of model parameters or labeled examples while controlling for the other variable by setting it to a constant.
Approach: They explore the dynamics of scaling up the number of model parameters versus the number labeled examples across a wide variety of tasks.
Outcome: The results show that scaling parameters yields performance improvements, while adding examples does not.
CORAL: Benchmarking Multi-turn Conversational Retrieval-Augmented Generation (2025.findings-naacl)

Copied to clipboard

Challenge: Existing research focuses on single-turn RAG, leaving a gap in addressing multi-turn conversations . a new benchmark is designed to assess RAG systems in realistic multi-turned conversations based on Wikipedia .
Approach: They propose a large-scale benchmark to assess RAG systems in multi-turn contexts . CORAL includes diverse information-seeking conversations automatically derived from Wikipedia . authors propose unified framework to standardize various conversational RAG methods .
Outcome: The proposed framework supports three core tasks of conversational RAG: passage retrieval, response generation, and citation labeling.
Extending Neural Generative Conversational Model using External Knowledge Sources (D18-1)

Copied to clipboard

Challenge: Existing generative dialogue models lack coherence and are content poor . however, current models lack the capacity to handle large unstructured knowledge sources.
Approach: They propose an architecture to incorporate unstructured knowledge sources to enhance the next utterance prediction in chit-chat type of generative dialogue models.
Outcome: The proposed architecture improves the next utterance prediction in chit-chat type of generative dialogue models by incorporating external knowledge from Wikipedia summaries and the NELL knowledge base.
CoXQL: A Dataset for Parsing Explanation Requests in Conversational XAI Systems (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing systems based on large language models (LLMs) are more precise and reliable in identifying users’ intentions, but the recognition of intents still presents a challenge in the case of ConvXAI, since little training data exist and the domain is highly specific.
Approach: They propose to use a dataset in the NLP domain for user intent recognition in ConvXAI to improve parsing performance.
Outcome: The proposed system outperforms existing methods and improves on existing ones.
Plan-Grounded Large Language Models for Dual Goal Conversational Settings (2024.eacl-long)

Copied to clipboard

Challenge: Existing studies show that LLMs can follow user instructions, but it is unclear how they can lead a plan-grounded conversation in mixed-initiative settings where instructions flow in both directions of the conversation.
Approach: They propose a dual-purpose mixed-initiative conversational setting where the LLM grounds the conversation on an arbitrary plan and seeks to satisfy both a procedural plan and user instructions.
Outcome: The proposed model achieves 2.1x improvement over a strong baseline and good generalization to unseen domains.
AirConcierge: Generating Task-Oriented Dialogue via Efficient Large-Scale Knowledge Retrieval (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing neural task-oriented dialogue systems cannot be encoded by memory networks, such as memory networks.
Approach: They propose an end-to-end trainable text-to SQL guided framework to learn a neural agent that interacts with KBs using the generated SQL queries.
Outcome: The proposed method significantly improves on the AirDialogue dataset, which contains the conversations of customers booking flight tickets from the agent.
Understanding User Utterances in a Dialog System for Caregiving (2020.lrec-1)

Copied to clipboard

Challenge: a dialog system that can monitor the health status of seniors has a huge potential for solving the labor shortage in the caregiving industry in aging societies.
Approach: They are developing a yes/no response classifier and an entailment recognizer to correctly interpret user utterances.
Outcome: The proposed system can correctly interpret user utterances and can monitor the health of seniors.
MKQA: A Linguistically Diverse Benchmark for Multilingual Open Domain Question Answering (2021.tacl-1)

Copied to clipboard

Challenge: Existing multilingual QA datasets lack linguistic diversity and comparable evaluation between languages.
Approach: They propose a multilingual question-answer evaluation set with 10k English queries and human translations of them into 25 additional languages and dialects.
Outcome: The proposed model is based on a multilingual knowledge questions and answers evaluation set with 26 languages.
Dialogue-AMR: Abstract Meaning Representation for Dialogue (2020.lrec-1)

Copied to clipboard

Challenge: Abstract Meaning Representation (AMR) does not capture the illocutionary force or speaker’s intended contribution in the broader dialogue context.
Approach: They propose a schema that enriches Abstract Meaning Representation (AMR) it provides a semantic representation for facilitating Natural Language Understanding (NLU) in dialogue systems.
Outcome: The proposed schema provides a semantic representation for facilitating Natural Language Understanding (NLU) in human-robot dialogue systems.
End-Task Oriented Textual Entailment via Deep Explorations of Inter-Sentence Interactions (P18-2)

Copied to clipboard

Challenge: Existing datasets for textual entailment (TE) have been used to study TE.
Approach: They propose a deep explorations of inter-sentence interactions for textual entailment task that uses a convolution to make important words in P and H play a dominant role in learnt representations.
Outcome: Experiments show that the pretrained DEISTE on SciTail gets 5% improvement over prior state of the art and that it generalizes well on RTE-5.
A Multi-Party Dialogue Ressource in French (2022.lrec-1)

Copied to clipboard

Challenge: a corpus of manual transcriptions of real-life, oral, spontaneous multi-party dialogues is available for French-speaking players of the board game Catan.
Approach: They propose to make available a corpus of manual transcriptions of real-life, oral, spontaneous multi-party dialogues between french-speaking players of the board game Catan.
Outcome: The proposed corpus is composed of long human-human interactions and can be used for dialogue studies in many fields.
Cross-lingual Intermediate Fine-tuning improves Dialogue State Tracking (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods to make multilingual systems expensive and tedious introduce pipeline of errors.
Approach: They propose to use pre-trained multilingual models to enhance the transfer learning process by intermediate fine-tuning of pretrained multi-lingual models.
Outcome: The proposed approach improves on the cross-lingual dialogue state tracking task with only 10% of the target language task data and zero-shot setup respectively.
Zero-shot Generalization in Dialog State Tracking through Generative Question Answering (2021.eacl-main)

Copied to clipboard

Challenge: Existing methods for Dialog State Tracking do not generalize well to new domains and unseen slots.
Approach: They propose an ontology-free framework that queries for unseen constraints and slots in multi-domain task-oriented dialogs using a conditional language model pre-trained on substantive English sentences.
Outcome: The proposed framework improves goal accuracy in zero-shot domain adaptation settings by up to 9% over the previous state-of-the-art on the MultiWOZ 2.1 dataset.
GCDST: A Graph-based and Copy-augmented Multi-domain Dialogue State Tracking (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to training DST on a single domain ignore information across domains.
Approach: They construct a dialogue state graph to transfer structured features among related domain-slot pairs across domains and encode the graph information of dialogue states by graph convolutional networks.
Outcome: The proposed model improves the performance of the multi-domain DST baseline with the absolute joint accuracy of 2.0% and 1.0% on the MultiWOZ 2.0 and 2.1 dialogue datasets.
SPAGHETTI: Open-Domain Question Answering from Heterogeneous Data Sources with Retrieval and Semantic Parsing (2024.findings-acl)

Copied to clipboard

Challenge: SPAGHETTI: Semantic Parsing Augmented Generation for Hybrid English information from Text Tables and Infoboxes is a hybrid question-answering pipeline .
Approach: They propose a hybrid question-answering pipeline that leverages knowledge from multiple knowledge sources.
Outcome: The proposed approach achieves state-of-the-art on the Compmix dataset with 56.5% exact match rate.
FLOW-BENCH: Towards Conversational Generation of Enterprise Workflows (2025.emnlp-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) can be used to convert natural language (NL) instructions into structured business process automation (BPA) process artifacts.
Approach: They propose to use large language models to convert natural language (NL) instructions into structured business process automation (BPA) process artifacts.
Outcome: The proposed model can be used to translate NL into Python and convert it into widely adopted business process definition languages.
Language Patterns and Behaviour of the Peer Supporters in Multilingual Healthcare Conversational Forums (2022.lrec-1)

Copied to clipboard

Challenge: a quantitative linguistic analysis of multilingual peer supporters in health-focused WhatsApp forums in Kenya is needed.
Approach: They conduct a quantitative linguistic analysis of the language usage patterns of multilingual peer supporters in two health-focused WhatsApp forums in Kenya.
Outcome: The proposed language analyzer can be used to analyze language usage patterns in two health-focused WhatsApp forums in Kenya.
Multi-Turn Response Selection for Chatbots with Deep Attention Matching Network (P18-1)

Copied to clipboard

Challenge: Existing models for matching dialogue responses rely on semantic and functional dependencies . a recent study only uses the last utterance in context for matching a reply .
Approach: They propose a model that matches a response with its multi-turn context using attention.
Outcome: The proposed model outperforms the state-of-the-art models on two large-scale multi-turn response selection tasks.
ChatR1: Reinforcement Learning for Conversational Reasoning and Retrieval Augmented Question Answering (2026.acl-long)

Copied to clipboard

Challenge: Unlike static ‘rewrite, retrieve, and generate’ pipelines, ChatR1 interleaves search and reasoning across turns, enabling exploratory and adaptive behaviors learned through RL.
Approach: They propose a reasoning framework based on reinforcement learning (RL) for conversational question answering that interleaves search and reasoning across turns and provides turn-level feedback.
Outcome: The proposed framework outperforms competing models on five CQA datasets, measured by different metrics (F1, BERTScore, and LLM-as-judge).
Effective QA-Driven Annotation of Predicate–Argument Relations Across Languages (2026.eacl-long)

Copied to clipboard

Challenge: Explicit representations of predicate-argument relations are a cornerstone of natural language understanding.
Approach: They propose a cross-linguistic projection approach that reuses an English QA-SRL parser within a constrained translation and word-alignment pipeline to automatically generate question-answer annotations aligned with target-language predicates.
Outcome: The proposed approach outperforms strong multilingual LLMs in Hebrew, Russian, and French.
The Context-Dependent Additive Recurrent Neural Net (N18-1)

Copied to clipboard

Challenge: Contextual sequence mapping is one of the fundamental problems in Natural Language Processing (NLP).
Approach: They propose a new family of Recurrent Neural Networks that address contextual sequence mapping . they propose to use contextual signals to control the flow of information .
Outcome: The proposed architecture outperforms existing methods on dialog problem and language model . the proposed architectures are based on a novel family of recurrent neural networks .
Expert Evaluation of a Spoken Dialogue System in a Clinical Operating Room (L18-1)

Copied to clipboard

Challenge: With the emergence of new technologies, the surgical working environment becomes increasingly complex and comprises many medical devices which have to be monitored and controlled.
Approach: They propose to use natural spoken language to control surgical operating rooms to reduce the amount of staff needed during a procedure.
Outcome: The proposed system can control the operating room using natural spoken language and is evaluated by experts in the field of minimally invasive surgery.
Interaction-Aware Topic Model for Microblog Conversations through Network Embedding and User Attention (C18-1)

Copied to clipboard

Challenge: Existing topic models ignore that one discusses diverse topics when dynamically interacting with different people.
Approach: They propose an Interaction-Aware Topic Model (IATM) for microblog conversations by integrating network embedding and user attention.
Outcome: The proposed model is based on three real-world microblog datasets.
STREAQ: Selective Tiered Routing for Effective and Affordable Contact Center Quality Assurance (2025.emnlp-industry)

Copied to clipboard

Challenge: Traditional manual QA cannot scale to growing volumes, while fully automated evaluation using large language models presents a cost-performance trade-off.
Approach: They propose a two-tier selective routing framework to intelligently route queries between cost-efficient and high-capability models.
Outcome: The proposed model reduces daily costs by 48% while preserving critical performance.
Put Chatbot into Its Interlocutor’s Shoes: New Framework to Learn Chatbot Responding with Intention (2021.naacl-main)

Copied to clipboard

Challenge: Currently, most work on improving the fluency and coherence of chatbots is focused on making them more human-like.
Approach: They propose a framework to train chatbots to possess human-like intentions by making them learn from interactive conversation.
Outcome: The proposed framework includes a guiding chatbot and an interlocutor model that plays the role of humans.
EVI: Multilingual Spoken Dialogue Tasks and Dataset for Knowledge-Based Enrolment, Verification, and Identification (2022.findings-naacl)

Copied to clipboard

Challenge: Knowledge-based authentication is crucial for task-oriented spoken dialogue systems that offer personalised and privacy-focused services . e-learning systems should be able to enrol, identify, and verify new and recurring users based on their personal information .
Approach: They propose to formalise three authentication tasks and their evaluation protocols . they propose to use a spoken multilingual dataset with 5,506 spoken dialogues .
Outcome: The proposed models set the first competitive benchmarks and set directions for future research.
ProMISe: A Proactive Multi-turn Dialogue Dataset for Information-seeking Intent Resolution (2024.findings-eacl)

Copied to clipboard

Challenge: Work done during internship at Amazon Alexa AI.
Approach: They propose to use iterative suggested question-answering conversation to improve the trade-off between satisfaction of the user’s intent and keeping the information exchange natural.
Outcome: The proposed proposed question-answering conversation improves the satisfaction of the user’s intent while keeping the information exchange natural and cognitive load of the interaction minimal on the users.
Appraisal Framework for Clinical Empathy: A Novel Application to Breaking Bad News Conversations (2024.lrec-main)

Copied to clipboard

Challenge: Empathy is essential in healthcare communication.
Approach: They propose an annotation approach that draws on well-established frameworks for clinical empathy and breaking bad news conversations for considering the dynamic dynamics of discourse relations.
Outcome: The proposed model can be used to train models to detect causal relations involving empathy, a feature of systems that can provide feedback to medical professionals in training.
Topic-Driven and Knowledge-Aware Transformer for Dialogue Emotion Detection (2021.acl-long)

Copied to clipboard

Challenge: Emotion detection in dialogues requires the identification of thematic topics underlying a conversation, commonsense knowledge, and the intricate transition patterns between affective states.
Approach: They propose a Topic-Driven Knowledge-Aware Transformer model that integrates topic representation and commonsense knowledge from ATOMIC for dialogue emotion detection.
Outcome: The proposed model outperforms state-of-the-art models on four dialogue datasets . it can detect topics which help distinguish emotion categories, the authors show .
CLASS: A Design Framework for Building Intelligent Tutoring Systems Based on Learning Science principles (2023.findings-emnlp)

Copied to clipboard

Challenge: CLASS empowers ITS with two key capabilities: first, it equips it with essential problem-solving strategies, and second, it facilitates natural language interactions, fostering engaging student-tutor conversations.
Approach: They propose a design framework called Conversational Learning with Analytical Step-by-Step Strategies (CLASS) that empowers ITS with two key capabilities: first, a carefully curated dataset and second, facilitating natural language interactions.
Outcome: The proposed framework empowers ITS with two key capabilities: first, it equips it with essential problem-solving strategies, and second, it facilitates natural language interactions, fostering engaging student-tutor conversations.
Towards a more Robust Evaluation for Conversational Question Answering (2021.acl-short)

Copied to clipboard

Challenge: Conversational Question Answering (CQA) is a new form of NLP . it uses conversation history to extract the answer of the current question.
Approach: They propose to use conversation history to evaluate models which can access the ground truth answers of previous turns at each turn of the conversation.
Outcome: The proposed evaluation protocol severely limits the effectiveness of the proposed models in fully autonomous chatbots and leads to unsuspected biases in their behavior.
Data Selection for Multi-turn Dialogue Instruction Tuning (2026.findings-acl)

Copied to clipboard

Challenge: Instruction-tuned language models often use noisy multi-turn dialogue datasets with topic drift, repetitive chitchat, and mismatched answer formats across turns.
Approach: They propose a dialogue-level framework that scores whole conversations rather than isolated turns.
Outcome: The proposed framework outperforms strong single-turn selectors, dialogue-level LLM scorers and heuristic baselines on three multi-turn benchmarks and an in-domain Banking test set.
Proactive User Information Acquisition via Chats on User-Favored Topics (2025.findings-emnlp)

Copied to clipboard

Challenge: PIA tasks require a system to acquire user information without making the user feel abrupt while engaging in a chat on a predefined topic.
Approach: They propose a task to acquire user's answers to predefined questions without making the user feel abrupt while engaging in a chat on a predefined topic.
Outcome: The proposed system outperforms LLMs prompted with task instructions in a dataset of 650 PIA chats and shows that it is reasonably accurate.
Knowledge-augmented Self-training of A Question Rewriter for Conversational Knowledge Base Question Answering (2022.findings-emnlp)

Copied to clipboard

Challenge: Recent rise of conversational applications has promoted the development of conversation KBQA (ConvKBQA).
Approach: They propose a framework to produce a full-fledged rewritten question based on conversation history and then reason the answer by existing single-turn KBQA models.
Outcome: The proposed framework produces a full-fledged rewritten question based on the conversation history and reasoned the answer by existing single-turn KBQA models.
DialCrowd 2.0: A Quality-Focused Dialog System Crowdsourcing Toolkit (2022.lrec-1)

Copied to clipboard

Challenge: DialCrowd 2.0 helps requesters obtain higher quality data from human intelligence tasks.
Approach: They propose to use DialCrowd 2.0 to help requesters obtain higher quality data . they aim to improve the way requesters present tasks and facilitate effective communication with workers.
Outcome: The proposed toolkit enables requesters to obtain higher quality data by presenting tasks more clearly and facilitating effective communication with workers.
Repo4QA: Answering Coding Questions via Dense Retrieval on GitHub Repositories (2022.coling-1)

Copied to clipboard

Challenge: Stack Overflow and GitHub are open source communities that are gaining popularity . developers need to raise programming questions in coding forums and navigate to GitHub repositories .
Approach: They propose a questionrepository matching task that bridges the gap between repositories and real-world coding questions.
Outcome: The proposed model outperforms state-of-the-art methods on coding questions and repositories . it can find suitable coding repositoriels and bridge the gap between them .
HyKnow: End-to-End Task-Oriented Dialog Modeling with Hybrid Knowledge Management (2021.findings-acl)

Copied to clipboard

Challenge: Task-oriented dialog systems typically manage structured knowledge to guide goal-oriented conversations.
Approach: They propose a TOD system with hybrid knowledge management, HyKnow, which extends the belief state to manage both structured and unstructured knowledge.
Outcome: The proposed model outperforms existing TOD systems in the evaluation of a multiWOZ dataset on unstructured knowledge with strong end-to-end performance.
Sentiment Adaptive End-to-End Dialog Systems (P18-1)

Copied to clipboard

Challenge: Existing methods to train dialog systems only consider semantic inputs and under-utilize other user information.
Approach: They propose to include user sentiment in the end-to-end learning framework to make dialog systems more user-adaptive and effective.
Outcome: The proposed system improves on a bus information search task with sentiment information.
Conversational QA Dataset Generation with Answer Revision (2022.coling-1)

Copied to clipboard

Challenge: Existing frameworks for conversational question-answer generation generate a large-scale dataset based on input passages.
Approach: They propose a conversational question-answer generation framework that extracts question-worthy phrases from passages and generates corresponding questions considering previous conversations.
Outcome: The proposed framework improves the quality of synthetic data and can be used for domain adaptation of conversational question answering.
Supervised and Unsupervised Transfer Learning for Question Answering (N18-1)

Copied to clipboard

Challenge: Several QA scenarios and datasets have been introduced over the past few years.
Approach: They conduct extensive experiments to investigate the transferability of knowledge from a source QA dataset to a target dataset using two QA models.
Outcome: The proposed model outperforms the previous best model on TOEFL listening comprehension test by 7% on target datasets.
CMQA: A Dataset of Conditional Question Answering with Multiple-Span Answers (2022.coling-1)

Copied to clipboard

Challenge: Existing QA datasets only contain unconditional and parallel answers . conditional question answering with hierarchical multi-span answers is challenging for the community to solve .
Approach: They propose a conditional question answering task with hierarchical multi-span answers . they propose CMQA, which contains conditional and hierarchic samples .
Outcome: The proposed task can be used to build more reliable and sophisticated QA systems.
Generating Information-Seeking Conversations from Unlabeled Documents (2022.emnlp-main)

Copied to clipboard

Challenge: a novel framework for conversational question answering from unlabeled documents has been proposed . a large-scale dataset of synthetic conversations is available for use in real-world applications .
Approach: They propose a framework for conversational question answering from unlabeled documents . they propose 'SimSeek' framework that simulates conversation from unlabelled documents based on two scenarios .
Outcome: The proposed framework achieves state-of-the-art performance on a recent CQA benchmark, QuAC.
A Survey on Asking Clarification Questions Datasets in Conversational Systems (2023.acl-long)

Copied to clipboard

Challenge: Existing studies on Asking Clarification Questions (ACQs) are incomparable due to inconsistent data, experimental setups and evaluation strategies.
Approach: They analyse the current research status on Asking Clarification Questions (ACQs) and propose a set of evaluation metrics and benchmarks for multiple ACQs-related tasks.
Outcome: The proposed techniques are compared with the available datasets and evaluated against benchmarks.
Towards End-to-End Open Conversational Machine Reading (2023.findings-eacl)

Copied to clipboard

Challenge: Existing approaches to the problem of open-retrieval conversational machine reading (OR-CMR) use two separate modules to approach the problem's two successive sub-tasks.
Approach: They propose to model OR-CMR as a unified text-to-text task in a fully end-to end style and propose to use a text-based approach to solve the problem.
Outcome: Experiments on the ShARC and OR-ShARC dataset show that the proposed framework can generalize to different backbone models.
AuraDial: A Large-Scale Human-Centric Dialogue Dataset for Chinese AI Psychological Counseling (2025.findings-emnlp)

Copied to clipboard

Challenge: AuraDial is a large-scale, human-centric dialogue dataset for Chinese AI psychological counseling .
Approach: They propose a large-scale human-centric dialogue dataset for Chinese AI psychological counseling . they propose rephrasing-based data generation methodology to foster more human-like responses .
Outcome: The proposed dataset outperforms other datasets in generating human-like responses.
Exploiting domain-slot related keywords description for Few-Shot Cross-Domain Dialogue State Tracking (2022.emnlp-main)

Copied to clipboard

Challenge: Existing frameworks for dialogue state tracking with domain-slot-value labels are expensive . current models are limited due to high cost of data annotation and lack of data in some domains .
Approach: They propose a framework based on domain-slot related description to tackle the challenge of few-shot cross-domain DST.
Outcome: The proposed framework outperforms existing methods on MultiWOZ and gains strong slot accuracy compared to existing models.
Enhancing Visual Dialog Questioner with Entity-based Strategy Learning and Augmented Guesser (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to build a visual dialog (VD) Questioner do not provide explicit guidance for questioner to generate visually related and informative questions.
Approach: They propose a Related entity enhanced Questioner that learns entity-based questioning strategy from human dialogs.
Outcome: The proposed approach achieves state-of-the-art performance on image-guessing task and question diversity.
MotivGraph-SoIQ: Integrating Motivational Knowledge Graphs and Socratic Dialogue for Enhanced LLM Ideation (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) have limitations in grounding ideas and mitigating confirmation bias during refinement.
Approach: They propose a framework that integrates a Motivational Knowledge Graph with a Q-Driven Socratic Ideator to enhance LLM ideation.
Outcome: The proposed framework enhances LLM ideation by integrating a Motivational Knowledge Graph with a Q-Driven Socratic Ideator.
Finding Diamonds in Conversation Haystacks: A Benchmark for Conversational Data Retrieval (2025.emnlp-industry)

Copied to clipboard

Challenge: Our work identifies unique challenges in conversational data retrieval . large language model-based systems operate through open-ended interactions without predefined specifications.
Approach: They propose a benchmark to evaluate systems that retrieve conversation data for product insights.
Outcome: The benchmark provides a reliable standard for measuring conversational data retrieval performance.
The APVA-TURBO Approach To Question Answering in Knowledge Base (C18-1)

Copied to clipboard

Challenge: Existing query languages for question answering over knowledge bases are not capable of processing queries presented in human language directly.
Approach: They advocate a new model architecture that includes a verification mechanism for checking the correctness of predicted relations.
Outcome: The proposed approach dramatically improves the question answering performance.
DIRECT: Direct and Indirect Responses in Conversational Text Corpus (2021.findings-emnlp)

Copied to clipboard

Challenge: Neural conversation models have been able to generate fluent responses through training on a dialogue corpus, but they lack the ability to reveal the implied intentions of users.
Approach: They propose to train neural conversation models on a dialogue corpus that provides pragmatic paraphrases to advance techniques for natural language understanding in dialogue systems.
Outcome: The proposed corpus provides 71,498 pairs of indirect–direct utterance pairs accompanied by a multi-turn dialogue history extracted from the MultiWoZ dataset.
KPQA: A Metric for Generative Question Answering Using Keyphrase Weights (2021.naacl-main)

Copied to clipboard

Challenge: Existing n-gram similarity metrics fail to discriminate the incorrect answers due to the free-form of the answer.
Approach: They propose a new metric that assigns different weights to each token via keyphrase prediction to judge the correctness of GenQA.
Outcome: The proposed metric has a significantly higher correlation with human judgments than existing metrics in various datasets.
VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs (2025.findings-naacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have greatly influenced the development of Large Multi-modal Video Models.
Approach: They propose a benchmark to assess the proficiency of Large Multi-modal Video Models (LMMs) in detecting and localizing anomalies and inconsistencies in videos.
Outcome: The proposed benchmark assesses the proficiency of Video-LMMs in detecting and localizing anomalies and inconsistencies in videos.
Asking Clarification Questions in Knowledge-Based Question Answering (D19-1)

Copied to clipboard

Challenge: Existing clarification datasets with limited annotated examples do not address ambiguous phenomena.
Approach: They propose a dataset that allows users to ask clarification questions using open-domain examples.
Outcome: The proposed model achieves better performance than strong baselines and provides new challenges.
GRICE: A Grammar-based Dataset for Recovering Implicature and Conversational rEasoning (2021.findings-acl)

Copied to clipboard

Challenge: a grammar-based dialogue dataset, GRICE, is designed to bring implicature into pragmatic reasoning in conversations . implicature recovery is a key component of open-ended dialogue reasoning .
Approach: They propose a grammar-based dialogue dataset to bring implicature into pragmatic reasoning . they use a hierarchical grammar model to generate the entire dataset .
Outcome: The proposed model shows a significant performance gap between baseline methods and human models . the model shows an overall performance boost in conversational reasoning .
CCQA: A New Web-Scale Question Answering Dataset for Model Pre-Training (2022.findings-naacl)

Copied to clipboard

Challenge: Existing approaches to answer open domain questions rely on unlabeled text or synthetically generated question-answer pairs.
Approach: They propose a large-scale open-domain question-answering dataset based on the Common Crawl project that can be used to in-domain pre-train popular language models.
Outcome: The proposed dataset achieves promising results in zero-shot, low resource and fine-tuned settings across multiple tasks, models and benchmarks.
Challenging Reading Comprehension on Daily Conversation: Passage Completion on Multiparty Dialog (N18-1)

Copied to clipboard

Challenge: Existing approaches to reading comprehension on multiparty dialogs have focused on children's stories or newswire.
Approach: They propose a new corpus and a robust deep learning architecture for a task in reading comprehension on multiparty dialog.
Outcome: The proposed model outperforms the state-of-the-art model on a different genre using bidirectional LSTM, showing a 13.0+% improvement for longer dialogs.
Learning to Predict Persona Information for Dialogue Personalization without Explicit Persona Description (2023.findings-acl)

Copied to clipboard

Challenge: Existing approaches to personalize dialogue agents rely on explicit persona descriptions during inference, which severely limits their application in real-world scenarios.
Approach: They propose a method that learns to predict persona information based on the dialogue history to personalize dialogue agents without relying on explicit persona descriptions during inference.
Outcome: The proposed method improves the consistency and engagingness of generated responses when conditioning on the predicted profile of the dialogue agent.
MSCTD: A Multimodal Sentiment Chat Translation Dataset (2022.acl-long)

Copied to clipboard

Challenge: Multimodal machine translation and textual chat translation have received considerable attention . however, little research has been devoted to multimodal machine translator in conversations .
Approach: They propose a task to generate more accurate translations with the help of dialogue history and visual context.
Outcome: The proposed task can generate more accurate translations with the help of dialogue history and visual context.
Learning to Learn Semantic Parsers from Natural Language Supervision (D18-1)

Copied to clipboard

Challenge: Existing logical forms require a user to be familiar with the underlying structure to learn a semantic parser.
Approach: They propose a method for training semantic parsers from natural language feedback . they use natural language inputs to parse feedback to leverage it as a form of supervision .
Outcome: The proposed algorithm learns a semantic parser from users’ corrections expressed in natural language.
You Only Need One Model for Open-domain Question Answering (2022.emnlp-main)

Copied to clipboard

Challenge: Recent approaches to Open-domain Question Answering use external knowledge bases, but have separate parameters and are weakly-coupled during training.
Approach: They propose to use a single question answering model trained end-to-end to retrieve external knowledge and rerank passages with a separate reranked model.
Outcome: The proposed model outperforms the previous state-of-the-art model by 1.0 and 0.7 exact match scores on the Natural Questions and TriviaQA open datasets.
Who Is Speaking to Whom? Learning to Identify Utterance Addressee in Multi-Party Conversations (D19-1)

Copied to clipboard

Challenge: In multi-party conversations, addressee information is not always explicit . researchers have spent great efforts to understand conversations between two participants, which is known as multi-part conversation.
Approach: They propose a who-to-whom model which models users and utterances in a conversation session jointly in an interactive way.
Outcome: The proposed model outperforms baseline models on the Ubuntu Multi-Party Conversation Corpus and shows consistent improvements.
Multi-User MultiWOZ: Task-Oriented Dialogues among Multiple Users (2023.findings-emnlp)

Copied to clipboard

Challenge: a dataset of task-oriented dialogues assume conversations between the agent and one user at a time . but multi-user task-orientated dialogues are richer, containing deliberation and deliberations . a novel task is proposed to rewrite a task-focused query that retains only task-relevant information .
Approach: They propose to rewrite a task-oriented chat between two users as a concise task-orientated query that retains only task-relevant information and is directly consumable by the dialogue system.
Outcome: The proposed method surpasses existing models on multi-user dialogues and generalizes to unseen domains.
Towards Improved Multi-Source Attribution for Long-Form Answer Generation (2024.naacl-long)

Copied to clipboard

Challenge: Current LLMs struggle with attribution for long-form answers which require reasoning over multiple evidence sources.
Approach: They propose to improve attribution capability of large language models for long-form answer generation to multiple sources with multiple citations per sentence.
Outcome: The proposed model improves on a wide range of attribution benchmark datasets on PolitiICite, a multi-source attribution dataset based on PolitIcite articles .
Knowledge-enhanced Mixed-initiative Dialogue System for Emotional Support Conversations (2023.acl-long)

Copied to clipboard

Challenge: Experimental results show the superiority of a mixed-initiative framework for emotional support conversation (ESC) ESC systems are emerging to provide prompt and convenient emotional support for helpseekers, including mental health support, counseling or motivational interviewing.
Approach: They propose a knowledge-enhanced mixed-initiative framework that retrieves actual case knowledge from a large-scale mental health knowledge graph for generating mixed-initiative responses.
Outcome: The proposed framework retrieves actual case knowledge from a large-scale mental health knowledge graph for generating mixed-initiative responses.
How to Make Neural Natural Language Generation as Reliable as Templates in Task-Oriented Dialogue (2020.emnlp-main)

Copied to clipboard

Challenge: Neural Natural Language Generation (NLG) systems are well known for their unreliability.
Approach: They propose a data augmentation approach which restricts the output of a neural network and guarantees reliability.
Outcome: The proposed approach scored 100% in semantic accuracy on the E2E NLG Challenge dataset, the same as a template system.
Interpretation of Natural Language Rules in Conversational Machine Reading (D18-1)

Copied to clipboard

Challenge: Existing work on question answering problems requires the reading of text because it contains a recipe to derive an answer together with the reader’s background knowledge.
Approach: They formalise a task and develop a crowd-sourcing strategy to collect 37k task instances based on real-world rules and crowd-generated questions and scenarios.
Outcome: The proposed task is based on 37k task instances based in real-world rules and crowd-generated questions and scenarios.
TRAVEL: Tag-Aware Conversational FAQ Retrieval via Reinforcement Learning (2023.emnlp-main)

Copied to clipboard

Challenge: Existing methods aim to fully utilize the dynamic conversation context to enhance the semantic association between the user query and FAQ questions, but they are limited by noise and e.g., users may click questions they don't like, leading to inaccurate semantics modeling.
Approach: They propose to introduce tags of FAQ questions to reduce noise in the conversation context and integrate them into a reinforcement learning framework to minimize the negative impact of irrelevant information.
Outcome: The proposed method can eliminate irrelevant information and minimize negative impact of irrelevant information in the dynamic conversation context.
Continual Dialogue State Tracking via Example-Guided Question Answering (2023.emnlp-main)

Copied to clipboard

Challenge: Dialogue systems are frequently updated to accommodate new services, but naively updating them by continually training with data for new services causes catastrophic forgetting.
Approach: They propose to reformulate dialogue state tracking (DST) as a bundle of example-guided question answering tasks to minimize the task shift between services.
Outcome: The proposed model achieves state-of-the-art performance on DST continual learning metrics without relying on any complex regularization or parameter expansion methods.
ACCENT: An Automatic Event Commonsense Evaluation Metric for Open-Domain Dialogue Systems (2023.acl-long)

Copied to clipboard

Challenge: evaluating commonsense in dialogue systems remains an open challenge . despite the success of open-domain dialogue systems, systems struggle to produce commonsensical responses as humans do.
Approach: They propose an event commonsense evaluation metric empowered by commonsensence knowledge bases.
Outcome: The proposed metric achieves higher correlations with human judgments than baselines.
A Qualitative Comparison of CoQA, SQuAD 2.0 and QuAC (N19-1)

Copied to clipboard

Challenge: In response to this development, there have been a flurry of new datasets for question answering.
Approach: They propose to use SQuAD 2.0, QuAC, and CoQA to provide question answering on textual data.
Outcome: The proposed datasets provide complementary coverage of the first two aspects, but weak coverage of third.
Slot Attention with Value Normalization for Multi-Domain Dialogue State Tracking (2020.emnlp-main)

Copied to clipboard

Challenge: Existing dialogue state tracking approaches rely on ontology already defined, where all slots and their possible values are given.
Approach: They propose a new architecture to exploit domain ontology by using Slot Attention and Value Normalization . they supplement the annotation of supporting span for MultiWOZ 2.1, which is the shortest span in utterances to support the labeled value.
Outcome: The proposed architecture exploits ontology and can convert supporting spans to values.
You Sound Like Someone Who Watches Drama Movies: Towards Predicting Movie Preferences from Conversational Interactions (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods for conversational recommendation include collaborative filtering, content-based filtering and user reviews.
Approach: They propose to map a conversational user to most similar external reviewers, whose preferences are known, and adapt collaborative filtering techniques to estimate the current user’s preferences for new movies.
Outcome: The proposed method can improve the accuracy of predicting user ratings for new movies by exploiting conversation content and external data.
Learning a Cost-Effective Annotation Policy for Question Answering (2020.emnlp-main)

Copied to clipboard

Challenge: State-of-the-art question answering systems require large amounts of training data for which labeling is time consuming and thus expensive.
Approach: They propose a framework for annotating QA datasets that entails learning a cost-effective annotation policy and a semi-supervised annotation scheme.
Outcome: The proposed approach can reduce up to 21.1% of the annotation cost compared with traditional methods . the proposed approach is based on a cost-effective annotation policy and semi-supervised annotation scheme .
Beyond Persuasion: Towards Conversational Recommender System with Credible Explanations (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing CRSs can be highly persuasive, but they can be deceptive and can damage the long-term trust between users and the CRS.
Approach: They propose a method to enhance the credibility of CRS’s explanations by using a set of credibility-aware persuasive strategies and a post-hoc self-reflection process.
Outcome: The proposed method enhances the credibility of CRS’s explanations and refines them via post-hoc self-reflection.
VGNMN: Video-grounded Neural Module Networks for Video-Grounded Dialogue Systems (2022.naacl-main)

Copied to clipboard

Challenge: Neural module networks (NMN) have been used in image-grounded tasks such as Visual Question Answering (VQA) however, very limited work on NMN has been studied in the video-ground dialogue tasks.
Approach: They propose to use video as the grounding feature in video-grounded dialogues to model the information retrieval process in videogrounded language tasks as a pipeline of neural modules.
Outcome: The proposed model can achieve promising performance on video-grounded dialogue and QA benchmarks.
Do not let the history haunt you: Mitigating Compounding Errors in Conversational Question Answering (2020.lrec-1)

Copied to clipboard

Challenge: Existing approaches employ human-written ground-truth answers for answering conversational questions at test time, but in a realistic scenario, the CoQA model will not have access to ground-Truth answers.
Approach: They propose a sampling strategy that dynamically selects between target answers and model predictions during training, closely simulating the situation at test time.
Outcome: The proposed sampling strategy closely simulates the situation at test time and significantly lowers the performance of CoQA systems.
SuperDialseg: A Large-scale Dataset for Supervised Dialogue Segmentation (2023.emnlp-main)

Copied to clipboard

Challenge: Empirical studies show that supervised learning is extremely effective in in-domain datasets and models trained on SuperDialseg can achieve good generalization ability on out-of-domain data.
Approach: They propose a supervised definition of dialogue segmentation points using document-grounded dialogues and a large-scale supervised dataset called SuperDialseg.
Outcome: The proposed model can achieve good generalization ability on out-of-domain data.
Answering Ambiguous Questions through Generative Evidence Fusion and Round-Trip Prediction (2021.acl-long)

Copied to clipboard

Challenge: Open-domain question answering is a task to answer questions using passages with diverse topics.
Approach: They propose a model that aggregates evidence from multiple passages to adaptively predict a single answer or a set of question-answer pairs for ambiguous questions.
Outcome: The proposed model achieves state-of-the-art performance on AmbigQA dataset and shows competitive performance on NQ-Open and TriviaQA.
Incorporating External Knowledge into Machine Reading for Generative Question Answering (D19-1)

Copied to clipboard

Challenge: Existing knowledge-aware QA models do not have commonsense and background knowledge to answer nontrivial questions.
Approach: They propose a new neural model which exploits external knowledge to generate answers in natural language for a given question with context.
Outcome: The proposed model improves answer quality over existing models without knowledge and knowledge-aware models, a study shows . state officials in Hawaii confirmed that president Barack Obama was born in the U.S.
One Agent To Rule Them All: Towards Multi-agent Conversational AI (2022.findings-acl)

Copied to clipboard

Challenge: Increasing volume of conversational agents (CAs) on the market has resulted in users being burdened with learning and adopting multiple agents to accomplish their tasks.
Approach: They propose a task BBAI: Black-Box Agent Integration that integrates multiple black-box CAs at scale.
Outcome: The proposed system outperforms existing benchmarks in the BBAI: Black-Box Agent Integration task.
MetaQA: Combining Expert Agents for Multi-Skill Question Answering (2023.eacl-main)

Copied to clipboard

Challenge: Recent explosion of question-answering datasets and models has increased interest in generalization of models across multiple domains and formats.
Approach: They propose to combine expert agents with a flexible and training-efficient architecture that considers questions, answer predictions, and answer-prediction confidence scores to select the best answer among a list of answer predictions.
Outcome: The proposed model outperforms previous multi-agent and multi-dataset approaches and is highly data-efficient to train and adaptable to any QA format.
Evaluating Theory of Mind in Question Answering (D18-1)

Copied to clipboard

Challenge: a dataset is proposed for question answering models with respect to their capacity to reason about beliefs.
Approach: They propose a dataset for evaluating question answering models with respect to their capacity to reason about beliefs.
Outcome: The proposed dataset is inspired by theory-of-mind experiments that examine whether children are able to reason about beliefs of others.
DialFact: A Benchmark for Fact-Checking in Dialogue (2022.acl-long)

Copied to clipboard

Challenge: Existing fact-checking models trained on non-dialogue data fail to perform well on this task.
Approach: They propose a task of fact-checking in dialogue to improve fact- checking performance . they propose to use an annotated conversational claim and Wikipedia snippets as evidence .
Outcome: The proposed task improves fact-checking performance in dialogue.
STL-CQA: Structure-based Transformers with Localization and Encoding for Chart Question Answering (2020.emnlp-main)

Copied to clipboard

Challenge: Chart Question Answering (CQA) is a task of answering natural language questions about visualisations in the chart image.
Approach: They propose a method for Chart Question Answering which improves the question/answering through sequential elements localization, question encoding and then, a structural transformer-based learning approach.
Outcome: The proposed method outperforms state-of-the-art methods on various chart Q/A datasets while outperforming even human baseline.
Beyond task success: A closer look at jointly learning to see, ask, and GuessWhat (N19-1)

Copied to clipboard

Challenge: Existing systems that address the abilities that need to be put to work during conversations are lacking in terms of visual grounding.
Approach: They propose a visually-grounded dialogue state encoder which integrates visual grounding with dialogue system components.
Outcome: The proposed system improves the GuessWhat?! game by combining guessing and asking questions with multi-task learning.
Knowledge-Driven Slot Constraints for Goal-Oriented Dialogue Systems (2021.naacl-main)

Copied to clipboard

Challenge: Traditional goal-oriented dialogue systems allow execution of validation rules as a post-processing step after slots have been filled which can lead to error accumulation.
Approach: They propose a task of constraint violation detection based on knowledge-driven slot constraints . they propose methods to integrate external knowledge into the system and compare it to traditional rule-based pipeline approach .
Outcome: The proposed task compares to the existing system and a rule-based pipeline.
VD-BERT: A Unified Vision and Dialog Transformer with BERT (2020.emnlp-main)

Copied to clipboard

Challenge: Prior work focused on attention mechanisms to model complex interactions in visual dialog . a new framework for visual dialog is based on pretrained BERT language models .
Approach: They propose a framework for a vision-dialog Transformer that leverages pretrained BERT language models for Visual Dialog tasks.
Outcome: The proposed framework achieves the top position on the visual dialog leaderboard without pretraining on external vision-language data.
Novel Slot Detection: A Benchmark for Discovering Unknown Slot Types in the Task-Oriented Dialogue System (2021.acl-long)

Copied to clipboard

Challenge: Existing slot filling models can only recognize pre-defined in-domain slot types from a limited slot set.
Approach: They introduce a task, Novel Slot Detection, in the task-oriented dialogue system.
Outcome: The proposed task is based on two public NSD datasets and proposes strong baselines . it aims to identify a sequence of tokens and extract semantic constituents from user queries .
Hello Again! LLM-powered Personalized Agent for Long-term Dialogue (2025.naacl-long)

Copied to clipboard

Challenge: Existing dialogue systems focus on brief single-session interactions, neglecting real-world needs for long-term companionship and personalized interactions.
Approach: They propose a model-agnostic framework for long-term dialogue agents . they use event summary and persona management to enable reasoning .
Outcome: The proposed framework incorporates three independently tunable modules dedicated to event perception, persona extraction, and response generation.
QANom: Question-Answer driven SRL for Nominalizations (2020.coling-main)

Copied to clipboard

Challenge: Traditionally, SRL annotations focus on verbal predicates, but other types of predicate are frequent in natural language.
Approach: They propose a semantic scheme for capturing predicate-argument relations for nominalizations, termed QANom, using crowdsourcing and QA-driven annotations.
Outcome: The proposed scheme outperforms existing annotations and is useful for downstream tasks.
ChatGPT Is a Knowledgeable but Inexperienced Solver: An Investigation of Commonsense Problem in Large Language Models (2024.lrec-main)

Copied to clipboard

Challenge: acquiring and representing commonsense in machines has posed a long-standing challenge (Li et al., 2021; Zhang e t al, 2022; Zhou e al. 2023) .
Approach: They use a commonsense-based LLM to evaluate ChatGPT's commonsensing abilities by analyzing 11 datasets and generating knowledge descriptions.
Outcome: The proposed model can achieve good QA accuracies while still struggling with certain domains of datasets.
Dialogue Graph Modeling for Conversational Machine Reading (2021.findings-acl)

Copied to clipboard

Challenge: Existing methods for conversational machine reading (CMR) are not effective for capturing multiple objects in complex interactive scenarios.
Approach: They propose a dialogue graph modeling framework that captures explicit and implicit interactions hidden in the rule documents and a model that asks clarification questions to the machine.
Outcome: The proposed model exceeds the milestone accuracy score of 80% on the ShARC benchmark and achieves new state-of-the-art by first exceeding the milestone precision score of 90%.
End-to-End Conversational Search for Online Shopping with Utterance Transfer (2021.emnlp-main)

Copied to clipboard

Challenge: a new study proposes a conversational search system that integrates product attributes and dialog with search . but it faces two real world challenges: imperfect product schema/knowledge and lack of training dialog data .
Approach: They propose an end-to-end conversational search system that integrates search with text . they propose an utterance transfer approach that generates dialogue utterations from other domains .
Outcome: The proposed system outperforms the best tested baseline in a conversational search dataset for online shopping.
Value-Agnostic Conversational Semantic Parsing (2021.acl-long)

Copied to clipboard

Challenge: Existing models rely on rich representations of dialogue history that include all previously generated components of the output.
Approach: They propose a model that abstracts over values to focus prediction on type- and function-level context.
Outcome: The proposed model outperforms baseline models by 7.3% and 10.6% on SMCalFlow and TreeDST datasets.
MMCoQA: Conversational Question Answering over Text, Tables, and Images (2022.acl-long)

Copied to clipboard

Challenge: Existing conversational QA systems only use a single knowledge source, e.g., paragraphs or knowledge graph, and assume it contains enough evidence to extract answers to users' questions.
Approach: They propose a task to answer users' questions with multimodal knowledge sources via multi-turn conversations using a multimodal dataset.
Outcome: The proposed task brings a series of research challenges, including but not limited to priority, consistency, and complementarity of multimodal knowledge.
Curating a Large-Scale Motivational Interviewing Dataset Using Peer Support Forums (2022.coling-1)

Copied to clipboard

Challenge: Existing therapeutic chatbots lack large-scale conversations between clients and trained counselors . prior work has found that social media platforms such as Reddit are used to vent distress and peers are seen to actively respond to such posts.
Approach: They propose to use peer support platforms to scrape conversational data from Reddit to determine whether counselors' responses align with real therapeutic conversations.
Outcome: The proposed method achieved 97% coverage out of 17.3K responses, meaning that out of 16.8K responses labeled with a moderate agreement.
Smoothing Dialogue States for Open Conversational Machine Reading (2021.emnlp-main)

Copied to clipboard

Challenge: Existing studies train independent or pipeline systems for the two subtasks but are trivial by using hard-label decisions to activate question generation.
Approach: They propose a method to smooth two dialogue states in one decoder and bridge decision making and question generation to provide a richer dialogue state reference.
Outcome: The proposed method achieves state-of-the-art on the OR-ShARC dataset.
A Mutual Information Maximization Approach for the Spurious Solution Problem in Weakly Supervised Question Answering (2021.acl-long)

Copied to clipboard

Challenge: Weakly supervised question answering usually has only final answers as supervision signals while correct solutions are not provided.
Approach: They propose to explicitly exploit the semantic correlations between question-answer pairs and predicted answers by maximizing mutual information between question and answer pairs.
Outcome: The proposed method significantly outperforms previous learning methods in terms of task performance and is more effective in training models to produce correct solutions.
Planning-Guided Tutoring with Assessment-Driven Memory for Pedagogical LLM Tutors (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to simulate tutor behaviors or preferences fail to sustain high-quality pedagogical conversations that provide explicit stepwise scaffolding and adapt to learners’ evolving cognitive states.
Approach: They propose a planning-guided tutoring framework with an assessment-driven memory for multi-turn math dialogue tutoring.
Outcome: Experiments on multi-turn math tutoring benchmarks show that ScaffoldLM significantly improves pedagogical tutoring quality over strong baselines.
What is wrong with you?: Leveraging User Sentiment for Automatic Dialog Evaluation (2022.findings-acl)

Copied to clipboard

Challenge: Existing metrics for dialog evaluation are trained on human annotations, which is cumbersome to collect.
Approach: They propose to use user sentiment and other information as proxy to measure the quality of previous dialogs.
Outcome: The proposed model is comparable to models trained on human annotated data.
Contextual Fine-to-Coarse Distillation for Coarse-grained Response Selection in Open-Domain Conversations (2022.acl-long)

Copied to clipboard

Challenge: Existing studies focus on coarse-grained response selection in retrieval-based dialogue systems.
Approach: They propose a Contextual Fine-to-Coarse (CFC) distilled model for coarse-grained response selection in open-domain conversations.
Outcome: The proposed model improves over baseline methods on two datasets based on the Reddit comments dump and Twitter corpus compared with baseline methods.
BaSCo: An Annotated Basque-Spanish Code-Switching Corpus for Natural Language Understanding (2022.lrec-1)

Copied to clipboard

Challenge: Basque-Spanish code-switching is a widespread phenomenon among bilingual speakers in the Basque Country.
Approach: They propose to use annotated utterances to train bilingual chatbots in Basque and Spanish to cover the phenomenon of code-switching.
Outcome: The proposed corpus is the first with annotated linguistic resources encompassing Basque-Spanish code-switching.
A Pre-training Strategy for Zero-Resource Response Selection in Knowledge-Grounded Conversations (2021.acl-long)

Copied to clipboard

Challenge: Existing methods to train retrieval-based dialogue systems rely on crowd-sourced data . however, it is difficult to collect large-scale dialogues that are grounded on background knowledge .
Approach: They propose to decompose training of knowledge-grounded response selection into three tasks . they propose to combine query-passage matching task with query-dialogue history matching task .
Outcome: Experimental results show that the proposed model can perform comparable to existing methods . the retrieval-based system can leverage background knowledge when conversing with humans .
BotsTalk: Machine-sourced Framework for Automatic Curation of Large-scale Multi-skill Dialogue Datasets (2022.emnlp-main)

Copied to clipboard

Challenge: a number of largescale datasets targeting a specific conversational skill have recently become available.
Approach: They propose a framework where multiple agents grounded to specific skills participate in a conversation to automatically annotate multi-skill dialogues.
Outcome: The proposed framework can be used to build open-domain chatbots with diverse communicative skills.
M2QA: Multi-domain Multilingual Question Answering (2024.findings-emnlp)

Copied to clipboard

Challenge: Language varies along several axes, most importantly, language instance and domain . lack of evaluation datasets prevents transfer of NLP systems to non-dominant languages .
Approach: They propose a multi-domain multilingual question answering benchmark to explore cross-lingual cross-domain performance of fine-tuned models and state-of-the-art LLMs.
Outcome: The proposed benchmark compared 13,500 SQuAD 2.0-style question-answer instances in German, Turkish, and Chinese for the domains of product reviews, news, and creative writing.
Question Answering in Climate Adaptation for Agriculture: Model Development and Evaluation with Expert Feedback (2025.findings-acl)

Copied to clipboard

Challenge: Existing domain-specific question answering systems have generative capabilities, but their ability to answer climate adaptation questions remains unclear.
Approach: They propose an iterative framework that enables LLMs to dynamically aggregate information from heterogeneous sources, such as climate literature and structured tabular climate data from climate model projections and historical observations.
Outcome: The proposed framework enables LLMs to dynamically aggregate information from heterogeneous sources, such as text from climate literature and structured tabular climate data from climate model projections and historical observations.
xDial-Eval: A Multilingual Open-Domain Dialogue Evaluation Benchmark (2023.findings-emnlp)

Copied to clipboard

Challenge: Currently, human evaluation is the most reliable way to holistically judge the quality of the dialogue.
Approach: They propose to use English dialogue evaluation metrics to generalize them to other languages.
Outcome: The proposed metrics outperform OpenAI’s ChatGPT in terms of average Pearson correlations over all datasets and languages.
A Co-Attentive Cross-Lingual Neural Model for Dialogue Breakdown Detection (2020.coling-main)

Copied to clipboard

Challenge: Existing models for dialogue breakdown detection do not focus on preventing dialogue breakdowns.
Approach: They propose a model that integrates a pretrained cross-lingual language model and a co-attention network for dialogue breakdown detection.
Outcome: The proposed model outperforms all previous approaches on evaluation metrics in Japanese and English tracks in Dialogue Breakdown Detection Challenge 4 .
Integrating User History into Heterogeneous Graph for Dialogue Act Recognition (2020.coling-main)

Copied to clipboard

Challenge: Existing models cannot fully recognize the specific expressions given by users due to the informality and diversity of natural language expressions.
Approach: They propose a Heterogeneous User History graph convolution network which utilizes the user’s historical answers grouped by DA labels as additional clues to recognize the DA label of utterances.
Outcome: The proposed model outperforms the state-of-the-art methods on two benchmark datasets and shows that it integrates user’s historical answers.
Global Readiness of Language Technology for Healthcare: What Would It Take to Combat the Next Pandemic? (2022.coling-1)

Copied to clipboard

Challenge: Language Technology (LT) has been used in the COVID-19 pandemic, but only in a handful of languages.
Approach: They propose to use conversational agents for information dissemination and basic diagnosis in 15 Asian and African languages with varying resource-availability to test their knowledge of LT.
Outcome: The proposed research confirms the pitiful state of LT even for languages with large speaker bases, such as Sinhala and Hausa, and identifies the gaps that could help prioritize research and investment strategies in LT for healthcare.
Evaluating Coherence in Dialogue Systems using Entailment (N19-1)

Copied to clipboard

Challenge: Evaluating open-domain dialogue systems is difficult due to the diversity of possible correct answers.
Approach: They propose a set of metrics for evaluating topic coherence using distributed sentence representations and calculable approximations of human judgment using conversational coherency.
Outcome: The proposed metrics can be used as a surrogate for human judgment based on conversational coherence on large-scale datasets and provide an unbiased estimate for the quality of the responses.
SQLWOZ: A Realistic Task-Oriented Dialogue Dataset with SQL-Based Dialogue State Representation for Complex User Requirements (2025.emnlp-main)

Copied to clipboard

Challenge: Existing TOD datasets present simplified interactions with simple slot-value style constraints and preferences.
Approach: They propose a novel TOD dataset that captures complex user requirements using SQL statements.
Outcome: The proposed dataset captures complex, real-world user requirements.
Generating Multi-Aspect Queries for Conversational Search (2026.eacl-long)

Copied to clipboard

Challenge: Conversational information seeking (CIS) systems aim to model the user’s information need within the conversational context and retrieve the relevant information.
Approach: They propose a multi-aspect query generation and retrieval framework which uses Large Language Models to break the user utterance into multiple queries.
Outcome: The proposed framework outperforms state-of-the-art query rewriting methods on six widely used CIS datasets and fine-tunes the model on MASQ yields significant improvements.
Dealing with Data Scarcity in Spoken Question Answering (2024.lrec-main)

Copied to clipboard

Challenge: erroneous automatic speech recognition transcriptions and data scarcity hinder spoken QA models . paper focuses on using limited annotated data to improve spoken qa performance .
Approach: They propose a framework for utilizing limited annotated data effectively to improve spoken QA performance.
Outcome: The proposed model produces question-answer pairs from unannotated data with 5.5% relative gain over the model trained with annotated datasets.
MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues (2026.findings-acl)

Copied to clipboard

Challenge: Existing evaluation benchmarks for Multimodal Large Language Models (MLLMs) focus on single-turn question answering, overlooking the complexity of multi-turn dialogues in real-world scenarios.
Approach: They propose a video understanding benchmark for MLLMs in multi-turn dialogues that assesses six core competencies that focus on perceptivity and interactivity.
Outcome: The MT-Video-Bench evaluates 1,000 multi-turn dialogues from diverse domains and reveals significant performance discrepancies and limitations in handling multi-turned video dialogues.
Multi-Modal Open-Domain Dialogue (2021.emnlp-main)

Copied to clipboard

Challenge: Recent work in open-domain conversational agents has demonstrated that significant improvements in humanness and user preference can be achieved via massive scaling in both pre-training data and model size.
Approach: They combine open-domain dialogue agents with vision models to investigate human preferences and humanness.
Outcome: The proposed model outperforms existing models in multi-modal dialogue while performing as well as its predecessor (text-only) BlenderBot.
Distilling ChatGPT for Explainable Automated Student Answer Assessment (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing automated student answer assessment models lack explainable and faithful feedback.
Approach: They propose a framework that leverages ChatGPT for student answer scoring and rationale generation.
Outcome: The proposed method improves the overall QWK score by 11% compared to ChatGPT.
Conversation Initiation by Diverse News Contents Introduction (N19-1)

Copied to clipboard

Challenge: Existing conversation systems assume that the user always initiates conversation and focus on how to respond to the given user’s utterance.
Approach: They propose to generate initial utterance by summarizing and chatting about news articles to avoid boredom by relying on boilerplate utterrances like greetings.
Outcome: The proposed model outperforms baseline models and based on information retrieval based and generation based models.
SIMMC 2.0: A Task-oriented Dialog Dataset for Immersive Multimodal Conversations (2021.emnlp-main)

Copied to clipboard

Challenge: Existing task-oriented dialog datasets do not situate the dialog in the user’s multimodal context.
Approach: They propose to use a dataset to study multimodal task-oriented dialogs in the shopping domain to situate them in the user’s multimodal context.
Outcome: The proposed dataset includes 11K task-oriented user->assistant dialogs (117K utterances) in the shopping domain, grounded in immersive and photo-realistic scenes.
Reasoning Like a Doctor: Improving Medical Dialogue Systems via Diagnostic Reasoning Process Alignment (2024.findings-acl)

Copied to clipboard

Challenge: Medical dialogue systems have attracted significant attention for their potential to act as medical assistants.
Approach: They propose a framework that emulates clinicians' diagnostic reasoning processes and aligns with clinician preferences through thought process modeling.
Outcome: The proposed framework generates appropriate responses that relies on abductive and deductive diagnostic reasoning analyses and aligns with clinician preferences through thought process modeling.
Deep Reinforcement Learning-based Dialogue Policy with Graph Convolutional Q-network (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for deep reinforcement learning lack the ability to learn the relationship between dialogue states and actions.
Approach: They propose a graph-structured dialogue policy framework for task-oriented dialogue systems that uses bipartite graphs to construct two different bipartites and generate user-related and knowledge-related subgraphs.
Outcome: The proposed framework significantly improves the effectiveness and stability of dialogue policies.
Conversational Semantic Parsing (2020.emnlp-main)

Copied to clipboard

Challenge: Structured representations for task-oriented assistant systems are limited due to the limitations of the representation.
Approach: They propose a semantic representation for task-oriented conversational systems that can represent co-reference and context carryover.
Outcome: The proposed model improves the best results on ATIS, SNIPS, TOP and DSTC2 by up to 5 points for slot-carryover.
In Prospect and Retrospect: Reflective Memory Management for Long-term Personalized Dialogue Agents (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches to long-term dialogue memory management fail to capture the natural semantic structure of conversations, leading to fragmented and incomplete representations.
Approach: They propose a mechanism that integrates forward- and backward-looking reflections into a personalized memory bank for effective future retrieval.
Outcome: The proposed mechanism outperforms state-of-the-art benchmarks on a long-term dialogue memory model.
Data Collection for Empirically Determining the Necessary Information for Smooth Handover in Dialogue (2022.lrec-1)

Copied to clipboard

Challenge: Despite advances in deep learning, dialogue systems struggle to achieve fully autonomous transactions with users.
Approach: They conducted an experiment in which two operators switched periodically while performing chat, consultation, and sales tasks in dialogue.
Outcome: The results show that adjacency pairs are useful for recording conversation history . key-value pairs are also useful when there are underlying tasks, such as consultation and sales .
Analysis of Implicit Conditions in Database Search Dialogues (L18-1)

Copied to clipboard

Challenge: Annotators annotated 50 database search dialogues with database field tags . 10% of the utterances included non-database-field information, authors say .
Approach: They propose to annotate database search dialogues on real estate and analyse their utterances for database queries.
Outcome: The proposed method can extract the implicit conditions from user utterances and construct queries.
An Information-Providing Closed-Domain Human-Agent Interaction Corpus (L18-1)

Copied to clipboard

Challenge: a human-agent interaction corpus is a corpus of conversations between a user and an embodied conversational agent operated by a wizard of oz . data collected to create a 'corpus' with unexpected situations, such as misunderstandings, false information, and interruptions.
Approach: They propose a public corpus for Human-Agent Interaction where the agent is controlled by a Wizard of Oz.
Outcome: The proposed corpus is based on 15 conversations between users and a wizard of Oz agent . the data are used to create a corpus with unexpected situations, such as misunderstandings, false information, and interruptions.
DialogQAE: N-to-N Question Answer Pair Extraction from Customer Service Chatlog (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing work on question-answer extraction fails to integrate incomplete utterances from dialog context for composite QA retrieval.
Approach: They propose a task where questions and corresponding answers might be separated across different utterances.
Outcome: The proposed methods perform well on 5 customer service datasets and set a benchmark for N-to-N DialogQAE with utterance and session level evaluation metrics.
Real-Time Open-Domain Question Answering with Dense-Sparse Phrase Index (P19-1)

Copied to clipboard

Challenge: Existing open-domain question answering models require multiple documents on-demand for every input query.
Approach: They propose query-agnostic indexable representations of document phrases that can drastically speed up open-domain question answering.
Outcome: The proposed model can be trained and deployed even in a single 4-GPU server.
Learning to Ask Conversational Questions by Optimizing Levenshtein Distance (2021.acl-long)

Copied to clipboard

Challenge: Existing methods for estimating maximum likelihood are limited by easily learned tokens . Existing systems that generate questions based on dialogue context are limited in their ability to learn tokens.
Approach: They propose a framework that optimizes the minimum Levenshtein distance through explicit editing actions.
Outcome: The proposed framework outperforms state-of-the-art methods on two benchmark datasets and generalizes well on unseen data.
DynaEval: Unifying Turn and Dialogue Level Evaluation (2021.acl-long)

Copied to clipboard

Challenge: Existing evaluation metrics focus on the turn-level quality of a dialogue . a unified framework that holistically considers the quality of the entire dialogue is needed .
Approach: They propose a unified automatic evaluation framework which holistically considers the quality of the entire dialogue.
Outcome: The proposed framework outperforms the state-of-the-art dialogue coherence model and correlates strongly with human judgements across multiple evaluation aspects at both turn and dialogue level.
Hierarchical Transformer for Task Oriented Dialog Systems (2021.naacl-main)

Copied to clipboard

Challenge: Existing models for dialog generation are challenging to train using the standard Seq2Seq models.
Approach: They propose a framework for Hierarchical Transformer Encoders that can be morphed into any hierarchical transformer by using specially designed attention masks and positional encodings.
Outcome: The proposed framework can be morphed into any hierarchical encoder, including HRED and HIBERT like models, by using specially designed attention masks and positional encodings.
PACIFIC: Towards Proactive Conversational Question Answering over Tabular and Textual Data in Finance (2022.emnlp-main)

Copied to clipboard

Challenge: Existing studies on financial question answering systems focus on passively responding to user queries.
Approach: They propose a new dataset to facilitate conversational question answering over hybrid contexts in finance . they propose PACIFIC to combine clarification question generation and CQA .
Outcome: The proposed method performs multi-task learning over all sub-tasks in PACIFIC . it incorporates a simple ensemble strategy to alleviate error propagation issue .
Reference production in human-computer interaction: Issues for Corpus-based Referring Expression Generation (L18-1)

Copied to clipboard

Challenge: Referring Expression Generation studies often use web-based data collection tasks without a particular addressee in mind.
Approach: They developed a parallel corpus of monologue and dialogue referring expressions and an annotated corpus to compare instances produced in both modes of communication.
Outcome: The results suggest that human reference production may be affected by the presence of a second (specific) human participant as the receiver of the communication in a number of ways.
Synthesizing Conversations from Unlabeled Documents using Automatic Response Segmentation (2024.findings-acl)

Copied to clipboard

Challenge: Several datasets have been developed for building conversational question answering systems.
Approach: They propose a robust dialog synthesising method that learns segmentation instead of using sentence boundaries.
Outcome: The proposed method achieves superior quality when compared to WikiDialog . it also improves performance across OR-QuAC benchmarks .
Learn to Resolve Conversational Dependency: A Consistency Training Framework for Conversational Question Answering (2021.acl-long)

Copied to clipboard

Challenge: Existing approaches do not explicitly train QA models on how to resolve conversational dependency, and thus these models are limited in understanding human dialogues.
Approach: They propose a framework that generates self-contained questions that can be understood without the conversation history and then trains a QA model with the pairs of original and self-constructed questions using a consistency-based regularizer.
Outcome: The proposed framework improves the models’ performance by up to 1.2 F1 on QuAC, and 5.2 F1 for CANARD, while addressing the limitations of the existing approaches.
PhotoChat: A Human-Human Dialogue Dataset With Photo Sharing Behavior For Joint Image-Text Modeling (2021.acl-long)

Copied to clipboard

Challenge: PhotoChat contains 12k dialogues, each of which is paired with a user photo that is shared during the conversation.
Approach: They propose to use PhotoChat to facilitate research on image-text modeling by combining a photo-sharing intent prediction task and a picture retrieval task to retrieve the most relevant photo according to the dialogue context.
Outcome: The proposed tasks achieve 10.4% recall@1 and 58.1% F1 scores, indicating that the proposed dataset presents interesting yet challenging real-world problems.
Q-TOD: A Query-driven Task-oriented Dialogue System (2022.emnlp-main)

Copied to clipboard

Challenge: Existing pipelined task-oriented dialogue systems have difficulties adapting to unseen domains . end-to-end systems are plagued by large-scale knowledge bases in practice .
Approach: They propose a query-driven task-oriented dialogue system that extracts dialogue context information into a natural language query.
Outcome: The proposed system outperforms strong baselines and establishes a new state-of-the-art performance on three publicly available task-oriented dialogue datasets.
Modeling User Satisfaction Dynamics in Dialogue via Hawkes Process (2023.acl-long)

Copied to clipboard

Challenge: Existing estimators measure performance by user satisfaction but ignore satisfaction dynamics across turns.
Approach: They propose to use user satisfaction estimation to estimate performance of dialogue systems by using an estimator to simulate users.
Outcome: The proposed estimator outperforms existing estimators on four benchmark dialogue datasets.
Single-dataset Experts for Multi-dataset Question Answering (2021.emnlp-main)

Copied to clipboard

Challenge: Prior work has focused on training one network on multiple datasets to build a model that performs well on all of the training datasets and generalizes and transfers better to new datasets.
Approach: They combine multiple reading comprehension datasets to build a multi-dataset question answering model with an ensemble of single-data set experts.
Outcome: The proposed model outperforms baseline models in in-distribution accuracy and generalization and transfer performance.
TRACE the Evidence: Constructing Knowledge-Grounded Reasoning Chains for Retrieval-Augmented Generation (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing retrievers are not perfect and often include irrelevant documents in the retrieved set.
Approach: They propose to construct knowledge-grounded reasoning chains from retrieved documents to integrate supporting evidence into RAG models.
Outcome: The proposed model achieves an average performance improvement of 14.03% on three multi-hop QA datasets.
MultiDoc2Dial: Modeling Dialogues Grounded in Multiple Documents (2021.emnlp-main)

Copied to clipboard

Challenge: Existing work treats document-grounded dialogue modeling as a machine reading comprehension task based on a single document or passage.
Approach: They propose a task and dataset for modeling goal-oriented dialogues grounded in multiple documents.
Outcome: The proposed task and dataset address realistic scenarios where goal-oriented dialogues involve multiple topics and hence are grounded on different documents.
Transformers to Learn Hierarchical Contexts in Multiparty Dialogue for Span-based Question Answering (2020.acl-main)

Copied to clipboard

Challenge: Existing approaches to embedding in multiparty dialogues are poor for span-based question answering (QA)
Approach: They propose a novel approach to transformers that learns hierarchical representations in multiparty dialogue.
Outcome: The proposed model improves on the FriendsQA dataset by 3.8% and 1.4% over the two state-of-the-art models.
Enhancing Dialogue Symptom Diagnosis with Global Attention and Symptom Graph (D19-1)

Copied to clipboard

Challenge: Existing studies on symptom diagnosis based on EHRs focus on the standard electronic medical records, but the dialogues between doctors and patients that contain more rich information are not well studied.
Approach: They propose to build a global attention mechanism to capture more symptom related information and build symptom graphs to model the associations between symptoms rather than treating each symptom independently.
Outcome: The proposed model achieves the state-of-the-art on the constructed dataset.
Generating Extractive Answers: Gated Recurrent Memory Reader for Conversational Question Answering (2023.findings-emnlp)

Copied to clipboard

Challenge: Conversational question answering (CQA) requires models to extract answers from given contents to answer follow-up questions according to conversation history.
Approach: They propose a novel architecture that integrates extractive MRC models into a generalized sequence-to-sequence framework.
Outcome: The proposed architecture can use less storage space and consider historical memory deeply and selectively.
Dual Hierarchical Dialogue Policy Learning for Legal Inquisitive Conversational Agents (2026.findings-acl)

Copied to clipboard

Challenge: Existing systems for conversational AI are user-driven, but in many real-world situations, they do not extract information to achieve its own objectives.
Approach: They propose an inquisitive conversational agent that learns when and how to ask probing questions . they also propose a framework for a conversational ICA specifically tailored to the court .
Outcome: The proposed method outperforms single-agent RL baselines on a U.S. Supreme Court dataset.
Conversing by Reading: Contentful Neural Conversation with On-demand Machine Reading (P19-1)

Copied to clipboard

Challenge: a new approach to contentful neural conversation is proposed . end-to-end models are effective in learning fluent responses, but their responses are often vacuous and uninformative.
Approach: They propose a model that provides the conversation model with relevant text on the fly as a source of external knowledge.
Outcome: The proposed model improves the informativeness and diversity of generated output compared to previous methods.
Towards Enhanced Immersion and Agency for LLM-based Interactive Drama (2025.acl-long)

Copied to clipboard

Challenge: Existing studies have focused on the role of immersion and agency in interactive drama.
Approach: They propose a playwriting-guided generation method that helps LLMs craft dramatic stories with substantially improved structures and narrative quality.
Outcome: The proposed method improves storytelling quality and immersion and agency, while allowing agents to refine their reactions to align with the player’s intentions.
Rationally Reappraising ATIS-based Dialogue Systems (P19-1)

Copied to clipboard

Challenge: Recent state-of-the-art neural models have obtained F1-scores near 98% on the task of slot filling.
Approach: They propose to fix annotation errors in ATIS and propose a rule-based grammar for slot filling that achieves a 95.82% F1 score.
Outcome: The proposed grammar achieves a 95.82% F1-score on the ATIS domain.
Into the Unknown Unknowns: Engaged Human Learning through Participation in Language Model Agent Conversations (2024.emnlp-main)

Copied to clipboard

Challenge: Recent advances in language models (LMs) and retrieval-augmented generation (RAG) have led to more capable chatbots and generative search engines.
Approach: They propose to emulate the educational scenario where children/students learn by listening to and participating in conversations of their parents/teachers by watching and steering the discourse among several LM agents.
Outcome: The proposed system outperforms baseline methods on discourse trace and report quality and is preferred by 70% of participants over a search engine and 78% over sabota.
ProMQA: Question Answering Dataset for Multimodal Procedural Activity Understanding (2025.naacl-long)

Copied to clipboard

Challenge: Existing studies typically provide traditional, but less practical evaluation testbeds for multimodal systems.
Approach: They propose a novel evaluation dataset, ProMQA, to measure the advancement of systems in application-oriented scenarios.
Outcome: The proposed evaluation dataset reveals a significant gap between human and competitive multimodal models.
Enhancing Extractive Question Answering in Multiparty Dialogues with Logical Inference Memory Network (2025.coling-main)

Copied to clipboard

Challenge: Existing models for multiparty dialogue question answering (QA) do not consider logical inference relations in multiparty dialogs, leading to suboptimal performance.
Approach: They propose a memory network with logical inference for extractive QA in multiparty dialogues.
Outcome: The proposed model achieves state-of-the-art on Molweni and FriendsQA benchmarks.
CREPE: Open-Domain Question Answering with False Presuppositions (2023.acl-long)

Copied to clipboard

Challenge: Existing question answering datasets assume all questions have well defined answers.
Approach: They propose a QA dataset containing a distribution of false presuppositions . they find that 25% of questions contain false presumptions .
Outcome: The proposed model finds that 25% of questions contain false presuppositions . the model can find presuffpositions moderately well, but struggle when predicting correctness .
Social Bias Benchmark for Generation: A Comparison of Generation and QA-Based Evaluations (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for assessing social bias in large language models (LLMs) do not capture nuanced and context-dependent nature of natural language generation.
Approach: They propose a Bias Benchmark for Generation (BBG) that evaluates social bias in long-form generation by having LLMs generate continuations of story prompts.
Outcome: The proposed benchmark is based on the English BBQ and Korean BBQ datasets and compares it with multiplechoice BBQ evaluation.
A Sentiment and Emotion Aware Multimodal Multiparty Humor Recognition in Multilingual Conversational Setting (2022.coling-1)

Copied to clipboard

Challenge: Humor is an essential aspect of daily conversation, and people try to provoke humor in their talks.
Approach: They propose a multitask framework that annotates Hindi utterances with sentiment and emotion classes.
Outcome: The proposed framework improves on the recently released Hindi Humor dataset . it takes sentiment and emotion into account to understand humor .
DemMA: Dementia Multi-Turn Dialogue Agent with Expert-Guided Reasoning and Action Simulation (2026.findings-acl)

Copied to clipboard

Challenge: Simulating dementia patients with large language models is challenging due to the need to model cognitive impairment, emotional dynamics, and nonverbal behaviors over long conversations.
Approach: They propose an expert-guided dementia dialogue agent for multi-turn patient simulation . they introduce a framework that trains a single LLM to jointly generate reasoning traces, patient utterances, and aligned behavioral actions .
Outcome: The proposed model outperforms baselines in persona fidelity, clinical validity, and educational effectiveness.
Learning to Match Representations is Better for End-to-End Task-Oriented Dialog System (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing systems for task-oriented dialogue lack belief states as supervisory signals.
Approach: They propose a method for knowledge retrieval driven by matching representations . they use a matching signal extractor to extract matching representation between contexts and entities .
Outcome: Experiments on three standard benchmarks show that the proposed method performs better than existing approaches.
Dialogue Collection for Recording the Process of Building Common Ground in a Collaborative Task (2022.lrec-1)

Copied to clipboard

Challenge: Existing studies on the process of building common ground have not been well conducted.
Approach: They propose a method for recording the process of building common ground through a dialogue by using the intermediate result of a task.
Outcome: The proposed method can record the building common ground process by using the intermediate result of a task and can be estimated quite accurately.
Q2: Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question Answering (2021.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation methods for factual consistency in knowledge-grounded dialogues are unreliable and limit their applicability.
Approach: They propose an automatic evaluation metric for factual consistency in knowledge-grounded dialogue using automatic question generation and question answering.
Outcome: The proposed evaluation metric consistently shows higher correlation with human judgements.
Strategy-level Entrainment of Dialogue System Users in a Creative Visual Reference Resolution Task (2022.lrec-1)

Copied to clipboard

Challenge: entrainment is a phenomenon in which interlocutors start speaking more similarly to each other.
Approach: They propose to use crowd-sourced data to study entrainment of users playing a creative reference resolution game with an autonomous dialogue system.
Outcome: The proposed system adapts the user's descriptive strategy to one that is simpler to parse for the natural language understanding unit without impinging on their creativity.
Knowledge-Aware Graph-Enhanced GPT-2 for Dialogue State Tracking (2021.emnlp-main)

Copied to clipboard

Challenge: Existing models for dialogue state tracking are based on Graph Attention Networks . if the relationship between slots and values is modelled explicitly, this can be improved .
Approach: They propose a model architecture that augments GPT-2 with Graph Attention Networks to allow sequential prediction of slot values.
Outcome: The proposed architecture improves performance against a strong GPT-2 baseline and with sparsely supervised training.
Learning When to Retrieve, What to Rewrite, and How to Respond in Conversational QA (2024.findings-emnlp)

Copied to clipboard

Challenge: Understanding users’ contextual search intent when generating responses is an understudied topic for conversational question answering (QA).
Approach: They propose a method that allows LLMs to decide when to retrieve in RAG settings given a conversational context.
Outcome: The proposed method improves on three conversational QA datasets and criticizes the quality of generated responses.
MCˆ2: Multi-perspective Convolutional Cube for Conversational Machine Reading Comprehension (P19-1)

Copied to clipboard

Challenge: Existing models combine previous questions for conversation understanding and only employ recurrent neural networks (RNN) for reasoning.
Approach: They propose a multi-perspective convolutional cube model that integrates 1D and 2D convolutions with recurrent neural networks (RNN) to understand context from different perspectives.
Outcome: The proposed model is based on the Conversational Question Answering (CoQA) dataset and achieves state-of-the-art results.
E-ConvRec: A Large-Scale Conversational Recommendation Dataset for E-Commerce Customer Service (2022.lrec-1)

Copied to clipboard

Challenge: Recent research has focused on developing conversational recommendation system (CRS), which provides valuable recommendations to users through conversations.
Approach: They construct an authentic Chinese dialogue dataset consisting of over 25k dialogues and 770k utterances, which contains user profile, product knowledge base, and multiple sequential real conversations between users and recommenders.
Outcome: The proposed dataset contains user profile, product knowledge base, and multiple sequential real conversations between users and recommenders.
Zero-Shot Dialogue State Tracking via Cross-Task Transfer (2021.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to training a dialogue state tracking model require extensive annotated dialogue data.
Approach: They propose to transfer cross-task knowledge from general question answering corpora to QA model that can handle zero-shot DST.
Outcome: The proposed model improves existing zero-shot and few-shot results on MultiWoz and shows better generalization ability in unseen domains.
SPORTSINTERVIEW: A Large-Scale Sports Interview Benchmark for Entity-centric Dialogues (2022.lrec-1)

Copied to clipboard

Challenge: Existing knowledge grounded dialogue datasets only contain external knowledge from one dimension, which limits the diversity of knowledge sources and may contain unwanted bias.
Approach: They propose to use two types of external knowledge sources as knowledge grounding in an interview dataset to model human dialogues.
Outcome: The proposed dataset contains 150K interviews and 34K interviewees . it is larger in size and has more than one dimension of external knowledge linking . however, the performance of the proposed models is far from humans .
EmoInHindi: A Multi-label Emotion and Intensity Annotated Dataset in Hindi for Emotion Recognition in Dialogues (2022.lrec-1)

Copied to clipboard

Challenge: Existing datasets for emotion recognition in dialogues are in English . existing datasets are limited to a few languages like Hindi .
Approach: They propose a large conversational dataset in Hindi for multi-label emotion and intensity recognition in conversations . they use a Wizard-of-Oz manner to annotate dialogues with 16 emotion labels .
Outcome: The proposed dataset contains 1,814 dialogues with 44,247 utterances in Hindi . it is based on a Wizard-of-Oz manner and can detect emotions in conversation .
CityEQA: A Hierarchical LLM Agent on Embodied Question Answering Benchmark in City Space (2025.emnlp-main)

Copied to clipboard

Challenge: Embodied Question Answering (EQA) tasks are primarily focused on indoor environments, leaving the complexities of urban settings unexplored.
Approach: They propose a task where an embodied agent answers open-vocabulary questions in dynamic city spaces.
Outcome: The proposed agent achieves 60.7% of human-level answering accuracy compared to baselines . the proposed agent outperforms existing agents in open-ended city spaces .
PCQPR: Proactive Conversational Question Planning with Reflection (2024.emnlp-main)

Copied to clipboard

Challenge: Current CQG methods focus on immediate context without strategic consideration of the specified conversational outcome.
Approach: They propose a method that uses a planning algorithm inspired by Monte Carlo Tree Search to generate contextually relevant questions.
Outcome: The proposed approach surpasses existing methods in e-learning and customer service fields . it generates contextually appropriate questions strategically devised to reach a specified outcome .
STYLE: Improving Domain Transferability of Asking Clarification Questions in Large Language Model Powered Conversational Agents (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods for addressing ambiguities in conversational search systems are one-size-fits-all and struggle to achieve effective domain transferability.
Approach: They propose a method to provide search engines with strategies regarding when to ask clarification questions in a post-hoc manner.
Outcome: The proposed method improves search performance 10% on four unseen domains.
HutCRS: Hierarchical User-Interest Tracking for Conversational Recommender System (2023.emnlp-main)

Copied to clipboard

Challenge: Existing CRSs assume that users like all attributes of the target item and dislike those unrelated to it, which can introduce bias in attribute-level feedback and impede the system’s ability to accurately identify the target items.
Approach: They propose a framework that allows users to explicitly acquire user preferences through natural language conversations by providing explicit answers (yes/no) for each attribute they require.
Outcome: The proposed framework portrays the conversation as a hierarchical interest tree that consists of two stages.
Multi-Faceted Self-Consistent Preference Alignment for Query Rewriting in Conversational Search (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to rewrite ambiguous queries ignore feedback from query rewriting, passage retrieval and response generation in the rewritten process.
Approach: They propose to construct self-consistent preference alignment data to generate more diverse rewritten queries.
Outcome: The proposed method is effective in both in- and out-of-distribution scenarios.
Training Turn-by-Turn Verifiers for Dialogue Tutoring Agents: The Curious Case of LLMs as Your Coding Tutors (2025.findings-acl)

Copied to clipboard

Challenge: Existing studies have focused on coding tutoring, but their capabilities in guiding users to solve complex tasks remain underexplored.
Approach: They propose a novel agent workflow, Trace-and-Verify, which combines knowledge tracing to estimate a student’s knowledge state and turn-by-turn verification to ensure effective guidance toward task completion.
Outcome: The proposed agent workflow achieves significantly higher success rates than existing tutoring agents.
Knowledge Transfer from Answer Ranking to Answer Generation (2022.emnlp-main)

Copied to clipboard

Challenge: Recent studies show that Question Answering (QA) based on Answer Sentence Selection (AS2) can be improved by generating an improved answer from the top-k ranked answer sentences.
Approach: They propose to train a GenQA model by transferring knowledge from a trained AS2 model . they use top ranked candidate as the generation target and next k top rated candidates as context .
Outcome: The proposed model outperforms existing models on public and industrial datasets.
End-to-End Task-Oriented Dialogue Systems Based on Schema (2023.findings-acl)

Copied to clipboard

Challenge: Existing approaches for task-oriented dialogue systems rely on a unified schema across domains, but we propose a schema-aware model for task oriented dialogues based on 'slots'
Approach: They propose a schema-aware end-to-end neural network model for handling task-oriented dialogues based on a dynamic set of slots within a unified schema.
Outcome: The proposed model performs better on a well-known dataset than baselines on 'schema-guided dialogue' systems.
GRAF: Graph Retrieval Augmented by Facts for Romanian Legal Multi-Choice Question Answering (2025.findings-acl)

Copied to clipboard

Challenge: Question answering systems have been used for various domains and languages.
Approach: They propose a novel approach for question answering (QA) that combines a dataset of Romanian legal questions with a CROL corpus of laws.
Outcome: The proposed approach achieves competitive results with generally accepted state-of-the-art methods and even exceeds them in most settings.
Product Question Answering in E-Commerce: A Survey (2023.acl-long)

Copied to clipboard

Challenge: Product question answering (PQA) aims to automatically provide instant responses to customer’s questions in E-commerce platforms.
Approach: They categorize PQA studies into four problem settings in terms of the form of provided answers.
Outcome: The proposed methods capture the unique challenges of product question answering (PQA) .
Generating Responses that Reflect Meta Information in User-Generated Question Answer Pairs (2020.lrec-1)

Copied to clipboard

Challenge: Existing approaches to realize consistent personalities require expensive data collection.
Approach: They propose to collect question-answer pairs for particular characters from online users . meta information such as emotion and intimacy was also collected .
Outcome: The proposed method can be used to train neural conversational models with high quality questions and meta information.
Task-Driven and Experience-Based Question Answering Corpus for In-Home Robot Application in the House3D Virtual Environment (2022.lrec-1)

Copied to clipboard

Challenge: Question answering is an important part of natural language processing (NLP)
Approach: They propose to use TEQA to investigate the ability of agent task experience understanding for the long-term household task.
Outcome: The proposed corpus aims to investigate the ability of task experience understanding of agents for the daily question answering scenario on the ALFRED dataset.
Reduce Human Labor On Evaluating Conversational Information Retrieval System: A Human-Machine Collaboration Approach (2023.emnlp-main)

Copied to clipboard

Challenge: Evaluating conversational information retrieval systems requires a significant amount of human labor for annotation.
Approach: They propose to use human annotation to calibrate evaluation results to eliminate evaluation biases.
Outcome: The proposed method consumes less than 1% of human labor and achieves a consistency rate of 95%-99% with human evaluation results.
AUGUST: an Automatic Generation Understudy for Synthesizing Conversational Recommendation Datasets (2023.findings-acl)

Copied to clipboard

Challenge: Existing work on conversational recommendation systems lacks high-quality data . existing datasets lack large-scale and high-level data based on human annotators .
Approach: They propose an automatic dataset synthesis approach that generates large-scale recommendation dialogues using structured graphs based on user-item information from the real world.
Outcome: The proposed approach can generate large-scale and high-quality recommendation dialogues . it exploits user preferences, knowledge graphs, and conversation ability from existing datasets based on real-world data .
Chat or Learn: a Data-Driven Robust Question-Answering System (2020.lrec-1)

Copied to clipboard

Challenge: QA systems tend to perform poorly at chitchat, while data-driven chatbots are typically user-friendly but not goal-oriented .
Approach: They propose to use a controller to perform dialogue act classification and feed user input either to a sequence-to-sequence chatbot or to QA systems.
Outcome: The proposed system is a spoken QA application for the Google Home smart speaker.
We Are What We Repeatedly Do: Inducing and Deploying Habitual Schemas in Persona-Based Responses (2023.emnlp-main)

Copied to clipboard

Challenge: a variety of personas can be elicited from large language models, but they are opaque and unpredictable.
Approach: They propose an approach to dialogue generation that retrieves relevant schemas to condition a large language model to generate persona-based responses.
Outcome: The proposed method captures habitual knowledge and generates persona-based responses from a large language model.
CONQRR: Conversational Query Rewriting for Retrieval with Reinforcement Learning (2022.emnlp-main)

Copied to clipboard

Challenge: Existing models for conversational question answering require specific retrievers to understand user questions.
Approach: They develop a query rewriting model CONQRR that rewrites a conversational question into a standalone question.
Outcome: The proposed model achieves state-of-the-art on an open-domain conversational question answering dataset and is effective for two different off-the shelf retrievers.
PECAN: LLM-Guided Dynamic Progress Control with Attention-Guided Hierarchical Weighted Graph for Long-Document QA (2025.findings-acl)

Copied to clipboard

Challenge: Long-document Question Answering (QA) challenges with large-scale text and long-distance dependencies.
Approach: They propose a method that leverages large language models to control retrieval process . they propose 'attention-based' retrieval methods that construct hierarchical graphs .
Outcome: The proposed method achieves LLM-level performance while maintaining computational complexity comparable to RAG methods.
PALM: Pre-training an Autoencoding&Autoregressive Language Model for Context-conditioned Generation (2020.emnlp-main)

Copied to clipboard

Challenge: Existing techniques for natural language understanding and generation use autoencoding and/or autoregressive objectives to train models.
Approach: They propose a self-supervised pre-training scheme that pre-trains an autoencoding and autoregressive language model on a large unlabeled corpus for generating new text conditioned on context.
Outcome: The proposed scheme achieves state-of-the-art results on a variety of language generation benchmarks covering generative question answering, abstractive summarization and conversational response generation.
Problem-Oriented Segmentation and Retrieval: Case Study on Tutoring Conversations (2024.findings-emnlp)

Copied to clipboard

Challenge: POSR is a task of breaking down conversations into segments and linking each segment to the relevant reference item.
Approach: They propose a task that breaks down conversations into segments and links each segment to the relevant reference item.
Outcome: The proposed method outperforms independent segmentation pipelines and large language models on joint metrics.
User Willingness-aware Sales Talk Dataset (2025.coling-main)

Copied to clipboard

Challenge: Despite the importance of user willingness, to the best of our knowledge, no previous study has addressed the development of automated sales talk dialogue systems that explicitly consider user willingness.
Approach: They developed a user willingness–aware sales talk collection by leveraging the ecological validity concept to elicit natural user willingness.
Outcome: The proposed system elicited user willingness at the utterance level from multiple perspectives and was able to improve the user's intent to purchase.
Adaptive Query Rewriting: Aligning Rewriters through Marginal Probability of Conversational Answers (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods to incorporate retriever’s preference during the training of query rewriting models rely on extensive annotations such as in-domain rewrites and/or relevant passage labels, limiting their generalization and adaptation capabilities.
Approach: They propose a framework for training query rewriting models with limited rewrite annotations from seed datasets and completely no passage label.
Outcome: The proposed approach decontexualizes conversational queries into self-contained questions suitable for off-the-shelf retrievers.
DisastQA: A Comprehensive Benchmark for Evaluating Question Answering in Disaster Management (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for question answering (QA) are lacking in a high-stakes environment.
Approach: They propose a rigorously verified benchmark of 3,000 expert-annotated questions . they propose 'keypoint-based evaluation protocol' emphasizing factual completeness over verbosity .
Outcome: Experiments with 20 models reveal substantial divergences from general-purpose models such as MMLU-Pro.
Evaluation Paradigms in Question Answering (2021.emnlp-main)

Copied to clipboard

Challenge: Despite substantial overlap, subtle but significant distinctions exert an outsize influence on research . one paradigm values creating more intelligent QA systems, the other paradigm values building QA system that appeals to users.
Approach: They propose to use the Cranfield and Manchester paradigms to describe research working towards building human-like, intelligent QA systems.
Outcome: The proposed paradigms are based on the findings of two recent studies on question answering (QA) the Cranfield paradigm is not new, but the Manchester paradigm is christened as the most eclectic in QA .
Question Answering as Programming for Solving Time-Sensitive Questions (2023.emnlp-main)

Copied to clipboard

Challenge: Recent studies show that Large Language Models (LLMs) have shown remarkable intelligence in question answering.
Approach: They propose to reframe the Question Answering task as Programming to overcome this limitation by leveraging LLMs' superior ability in understanding both natural language and programming language.
Outcome: The proposed approach improves on time-sensitive question answering datasets by 14.5% over baselines.
QUARTZ: QA-based Unsupervised Abstractive Refinement for Task-oriented Dialogue Summarization (2025.findings-emnlp)

Copied to clipboard

Challenge: a framework for task-oriented utility-based dialogue summarization is proposed . QUARTZ is a tool for task summarizing dialogues, but its outputs lack task-specific focus.
Approach: They propose a framework for task-oriented utility-based dialogue summarization . QUARTZ generates summaries and question-answer pairs from a dialogue in a zero-shot manner .
Outcome: The proposed framework achieves competitive results in zero-shot settings, rivaling fully-supervised State-of-the-Art methods.
Introducing CQuAE : A New French Contextualised Question-Answering Corpus for the Education Domain (2024.lrec-main)

Copied to clipboard

Challenge: a new question answering corpus in french is designed to educational domain . we propose more complex questions and can justify the answers on validated material .
Approach: They propose a question answering corpus in French designed to educational domain . they propose to propose more complex questions and justify answers on validated material .
Outcome: The proposed question answering corpus is designed to be useful in educational domain . it proposes more complex questions and can justify answers on validated material . the proposed corpus could be used in the education domain, but it's not yet ready for use .
Downstream Trade-offs of a Family of Text Watermarks (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) can generate humanlike responses to a variety of requests like writing emails, translating or summarizing content.
Approach: They evaluate the performance of large language models (LLMs) watermarked using three different strategies over a diverse suite of tasks including those cast as k-class classification (CLS), multiple choice question answering (MCQ), short-form generation (e.g., open-ended question answering) and long-form generator (eg. translation)
Outcome: The proposed models can cause significant drops in their effectiveness across a variety of tasks including CLS, MCQ, short-form generation and translation tasks.
Beyond Chunking: Discourse-Aware Hierarchical Retrieval for Long Document Question Answering (2026.acl-long)

Copied to clipboard

Challenge: Existing long document question answering systems process texts as flat sequences or use heuristic chunking, which overlooks the discourse structures that guide human comprehension.
Approach: They propose a discourse-aware hierarchical framework that leverages rhetorical structure theory for long document question answering.
Outcome: The proposed framework exhibits strong robustness across diverse document types and linguistic settings.
Minimal Yet Big Impact: How AI Agent Back-channeling Enhances Conversational Engagement through Conversation Persistence and Context Richness (2024.findings-emnlp)

Copied to clipboard

Challenge: Increasing use of AI agents in conversational services highlights the importance of back-channeling (BC) as an active listening strategy to enhance conversational engagement.
Approach: They conducted an experiment with 55 participants to evaluate conversational engagement using both quantitative and qualitative metrics.
Outcome: The results show that the Todak_BC and TodAK_NoBC groups have significantly higher conversational engagement than the Todask_NoB.
Bridging The Gap: Entailment Fused-T5 for Open-retrieval Conversational Machine Reading Comprehension (2023.acl-long)

Copied to clipboard

Challenge: Open-retrieval conversational machine reading comprehension (OCMRC) simulates real-life conversation scenes.
Approach: They propose a one-stage end-to-end framework to bridge the information gap between decision-making and question generation in a global understanding manner.
Outcome: The proposed framework achieves new state-of-the-art performance on the OR-ShARC benchmark.
Know the Known and the Unknown: Reasonable Answer Generation with Knowledge-Informed Citations (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches focus on generating multi-level citations linked to specific references, making it verifiable and trustworthy.
Approach: They propose a new data construction pipeline and a benchmark to improve citation granularity and awareness of unknown information.
Outcome: The proposed model improves on the existing benchmark and data construction pipeline and provides citation granularity and awareness of unknown information.
Measuring Retrieval Complexity in Question Answering Systems (2024.findings-acl)

Copied to clipboard

Challenge: a new metric, retrieval complexity (RC), measures the difficulty of answering questions.
Approach: They propose a retrieval complexity metric conditioned on the completeness of retrieved documents . they propose an unsupervised pipeline to measure RC given an arbitrary retrieval system .
Outcome: The proposed pipeline measures RC more accurately than alternative estimators on six challenging QA benchmarks.
FinLFQA: Evaluating Attributed Text Generation of LLMs in Financial Long-Form Question Answering (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing benchmarks focus on simple attribution that retrieves textual evidence as references.
Approach: They propose a benchmark to evaluate the ability of large language models to generate reliable attributions.
Outcome: The proposed benchmark evaluates the ability of LLMs to generate long-form answers with reliable and nuanced attributions.
LLaMA-Omni 2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis (2025.acl-long)

Copied to clipboard

Challenge: LLaMA-Omni 2 is a series of speech language models (SpeechLMs) based on large language models.
Approach: They introduce a series of speech language models capable of real-time speech interaction . LLaMA-Omni 2 trains on 200K multi-turn speech dialogue samples .
Outcome: The proposed speech language models surpass state-of-the-art models on spoken question answering and speech instruction.
Thread: A Logic-Based Data Organization Paradigm for How-To Question Answering with Retrieval Augmented Generation (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in retrieval-augmented generation (RAG) have substantially improved question-answering systems, particularly for factoid ‘5Ws’ questions.
Approach: They propose a data organization paradigm where large language models transform documents into more structured and loosely interconnected LUs.
Outcome: Experiments in open-domain and industrial settings show that the proposed paradigm outperforms existing paradigms and shows high adaptability across diverse document formats.
Recursive Question Understanding for Complex Question Answering over Heterogeneous Personal Data (2025.findings-acl)

Copied to clipboard

Challenge: a novel method for question answering over mixed sources, like text and tables, has been developed for question-answering . personal information is a prominent case of such heterogeneous data, such as calendar entries, workout statistics, shopping records, streaming history, and more.
Approach: They propose a method that creates an executable operator tree for a given question . they use recursive decomposition to decompose a question into an operator tree .
Outcome: The proposed method outperforms methods based on verbalization or translation . it can be executed on user devices and yields a traceable answer .
StorySparkQA: Expert-Annotated QA Pairs with Real-World Knowledge for Children’s Story-Based Learning (2024.emnlp-main)

Copied to clipboard

Challenge: Existing story reading systems fail to capture the nuances of how education experts think when conducting interactive story reading activities.
Approach: They propose to use existing question-answering (QA) datasets to capture experts' annotations and thinking process to construct a story-based annotation framework.
Outcome: The proposed framework captures experts’ annotations and thinking process and can be used to generate 5, 868 expert-annotated QA pairs with real-world knowledge.
MP2D: An Automated Topic Shift Dialogue Generation Framework Leveraging Knowledge Graphs (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods to manage topic shifts within on-topic dialogues are limited in their ability to generate training datasets.
Approach: They propose a data generation framework that automatically generates conversational question-answering datasets with natural topic transitions by leveraging relationships between entities in a knowledge graph.
Outcome: The proposed framework generates conversational question-answering datasets with natural topic transitions and proves its effectiveness in generating dialogues with topic shifts.
Graph vs. Sequence: An Empirical Study on Knowledge Forms for Knowledge-Grounded Dialogue (2023.emnlp-main)

Copied to clipboard

Challenge: Knowledge-grounded dialogue systems can generate informative responses based on dialogue history and external knowledge source.
Approach: They conduct a thorough experiment to determine the optimal knowledge form, mutual effects between knowl- edge and model selection, and the few-shot performance of knowledge.
Outcome: The proposed method combines knowledge-grounded dialogue with human-generated dialogues to generate informative and meaningful responses.
Tailoring Diagnostic Modeling to Individual Learners: Personalized Distractor Generation via MCTS-Guided Reasoning Reconstruction (2026.acl-long)

Copied to clipboard

Challenge: Current distractor generation methods produce shared distractors for all students, ignoring individual variations in reasoning, which limits their diagnostic effectiveness.
Approach: They propose a method which tailors distractors to each student’s specific cognitive flaws, inferred from their past question-answering (QA) history.
Outcome: The proposed framework outperforms existing methods in generating plausible distractors and adapts to group-level settings.
A Survey of Ontology Expansion for Conversational Understanding (2024.emnlp-main)

Copied to clipboard

Challenge: Current methods for conversational understanding rely on static ontologies, limiting their ability to handle new and unforeseen user needs.
Approach: They propose to review the state-of-the-art techniques in OnExp for conversational understanding and highlight emerging frontiers . they categorize existing literature into three main areas: (1) New Intent Discovery, (2) New Slot-Value Discovery, and (3) Joint OnExp.
Outcome: The proposed methods highlight several emerging frontiers in OnExp to improve agent performance in real-world scenarios and discuss their corresponding challenges.
PK-ICR: Persona-Knowledge Interactive Multi-Context Retrieval for Grounded Dialogue (2023.emnlp-main)

Copied to clipboard

Challenge: Identifying relevant persona or knowledge for conversational systems is difficult, but recent work has shown that it is more realistic to optimize for concrete persona.
Approach: They propose a persona-knowledge dual context retrieval method that utilizes all dialogue contexts simultaneously.
Outcome: The proposed method performs zero-shot top-1 knowledge retrieval and precise persona scoring.
Comparative Analysis of the Intrinsic Metrics for Tokenizers and their effect on Downstream Tasks for Hindi and Marathi (2026.acl-long)

Copied to clipboard

Challenge: Various studies have shown that the performance of language models is poor in non-English or non-European languages.
Approach: They propose a grapheme cluster tokenizer which shows better performance than other popular tokenizers.
Outcome: The proposed tokenizers show better or competitiveness on question-answering tasks . the proposed tokenization model is highly correlated to the performance of other tokenizer models .
Do LVLMs Know What They Know? A Systematic Study of Knowledge Boundary Perception in LVLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Vision-Language Models (LVLMs) demonstrate strong visual question answering (VQA) capabilities but are shown to hallucinate.
Approach: They propose three confidence-based methods to enhance LVLMs' perception . they propose probabilistic and consistency-based signals are more reliable indicators .
Outcome: Experiments on three LVLMs across three VQA datasets show that LVLs possess a reasonable perception level but there is room for improvement.
MMRC: A Large-Scale Benchmark for Understanding Multimodal Large Language Model in Real-World Conversation (2025.acl-long)

Copied to clipboard

Challenge: Existing multimodal large language models lack the ability to memorize, recall, and reason in sustained interactions.
Approach: They propose a multimodal real-world conversation benchmark for evaluating open-ended abilities of multimodal large language models.
Outcome: The proposed benchmarks show that the models perform better in open-ended conversations.
One Planner To Guide Them All ! Learning Adaptive Conversational Planners for Goal-oriented Dialogues (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for goal-oriented dialogues involve training separate models for specific combinations of objectives, leading to computational and scalability issues.
Approach: They propose a new dialogue policy method that can adapt to varying objective preferences at inference time without retraining.
Outcome: The proposed method can adapt to varying objective preferences at inference time without retraining.
Reimagining Intent Prediction: Insights from Graph-Based Dialogue Modeling and Sentence Encoders (2024.lrec-main)

Copied to clipboard

Challenge: Existing approaches to intent prediction are limited in highly specialized fields, such as closed-domain dialogue systems, where context comprehension is of paramount importance.
Approach: They propose a method that uses scenario dialog graphs to model dialogues as sequences of transitions between intents, representing distinct goals or requests.
Outcome: The proposed method significantly advances the field of dialogue systems, providing valuable insights into the effectiveness and potential limitations of the proposed approaches.
LongRAG: A Dual-Perspective Retrieval-Augmented Generation Paradigm for Long-Context Question Answering (2024.emnlp-main)

Copied to clipboard

Challenge: Existing long-context Large Language Models (LLMs) struggle with the “lost in the middle” issue.
Approach: They propose a general, dual-perspective, and robust LLM-based RAG system paradigm for LCQA to enhance RAG’s understanding of complex long-context knowledge.
Outcome: The proposed system outperforms long-context LLMs, advanced RAG, and vanilla RAG on three multi-hop datasets.
TALON: A Multi-Agent Framework for Long-Table Exploration and Question Answering (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to query-relevant content retrieval fail to retrieve contextually relevant data.
Approach: They propose a multi-agent framework for table question answering over long tables . TALON features a planning agent that iteratively invokes a tool agent to access tabular data .
Outcome: The proposed framework achieves average accuracy improvements of 7.5% and 12.0% across all language models.
MAviS: A Multimodal Conversational Assistant For Avian Species (2025.emnlp-main)

Copied to clipboard

Challenge: Existing multimodal large language models face challenges when it comes to specialized topics like avian species.
Approach: They propose a large-scale multimodal avian species dataset that integrates image, audio, and text modalities for over 1,000 bird species.
Outcome: The proposed model outperforms the baseline MiniCPM-o-2.6 by a large margin.
NewsInterview: a Dataset and a Playground to Evaluate LLMs’ Grounding Gap via Informational Interviews (2025.acl-long)

Copied to clipboard

Challenge: Existing large datasets (1k-10k transcripts) are generated via crowdsourcing and are inherently unnatural.
Approach: They curate a dataset of 40,000 two-person informational interviews from NPR and CNN . they find that LLMs are significantly less likely than human interviewers to use acknowledgements and pivot to higher-level questions.
Outcome: The proposed model is based on 40,000 interviews with journalists and CNN .
Taking Notes Brings Focus? Towards Multi-Turn Multimodal Dialogue Learning (2025.emnlp-main)

Copied to clipboard

Challenge: Existing multimodal large language models are trained on single-turn vision question-answering tasks, which do not accurately reflect real-world human conversations.
Approach: They propose a large-scale multi-turn multimodal dialogue dataset that uses rules and GPT assistance to generate a multi-turned multimodal dialog dataset.
Outcome: The proposed dataset is a strong benchmark for multi-turn multimodal dialogue learning . it features complex dialogues with contextual dependencies that force models to track, ground, and recall information across multiple turns and disparate visual regions.
ChatAnime: Towards User-Centered Emotional Support in LLM-based Virtual Character Chat (2026.acl-long)

Copied to clipboard

Challenge: Existing research focuses on character consistency in fictional or game-based scenarios . ESRP framework is designed to align role-playing with real-world user scenarios based on emotional needs.
Approach: They propose a framework to align role-playing with real-world user scenarios and emotional needs.
Outcome: The proposed framework aligns role-playing with real-world user scenarios and emotional needs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations